Why most hospital AI pilots never reach production
I have sat through a lot of pilot review meetings. The pattern is consistent enough to be funny. The deck says the pilot met its objectives. The clinicians who participated say kind things. Somebody asks what it would take to roll this out to the other units, and the room goes quiet, because the honest answer is eleven things nobody has started and four decisions nobody is empowered to make.
The gap between a successful pilot and a production system in a hospital group is not a scaling problem. It is a set of commitments — an owner, a run-team, an integration path, a budget line, a change programme — that a pilot is specifically designed to avoid making. That is what makes pilots attractive. You get to look like you are moving without anyone signing anything.
So most groups end up with a portfolio of successful pilots and nothing in production. I have contributed to that portfolio. Here is where it actually breaks, and which decisions I now insist on making before the pilot starts rather than after.
The pilot succeeds because it was designed to
A pilot runs in one unit, usually the one with the best IT and the most cooperative unit head. It runs on a data extract rather than a live interface. It is supported personally by whoever is running it, who answers their phone at nine at night because it is their project. It uses a vendor’s team working at a level of attention no commercial contract will sustain. And its success criteria were written by the people who wanted it to succeed.
None of those conditions survive contact with unit four. The point is not that pilots are dishonest — they are useful for learning whether clinicians will accept something. The point is that a pilot tests the idea and production tests the organisation, and passing the first tells you almost nothing about the second.
The tell is simple. If the answer to “who supports this at nine at night in production” is a name rather than a function, you are in a pilot regardless of what the status report calls it.
Nobody owns it
This is the first and largest failure. In a multi-unit group, a production AI system needs an owner who holds three things together: the clinical standard it must meet, the operational process it sits inside, and the technology that runs it. Most groups have three different people holding one each, and no one accountable for the whole.
The digital or IT function cannot own a clinical process. The medical administration cannot own a technology platform’s uptime. The unit cannot own something that spans units. So the system ends up owned by a steering committee, which means owned by nobody, which means that when the interface breaks on a Sunday nothing happens until Monday and by Wednesday the unit has gone back to the old process permanently.
What has worked for me: a named business owner who runs the process — contact centre head for an enquiry agent, medical administration for a documentation system, the diagnostics lead for anything touching reports — with the technology function accountable to them on a defined service level, and a clinical authority who owns the standard and the never-do list but not the operation. Three roles, one decision-maker, written down before the pilot.
There is no run-team
Pilots are run by project people. Production needs an operations function, and in hospital groups it usually does not exist for anything that is not the HIS.
What production actually requires, and what almost no pilot has:
- A first line that knows the system well enough to answer a nurse’s question at eleven at night and triage a genuine fault from a misunderstanding. The contact centre becomes this by default, with no training and no mandate, and resents it.
- A second line with access to logs and a runbook for the five failures you already know about — interface down, model timing out, schedule master out of sync, language detection misfiring, queue backing up.
- Somebody whose job is to watch the quality metrics weekly. Not accuracy in the abstract — escalation rate, edit burden, handover pickup time, the unresolved queue. These drift, and drift is invisible without an owner.
- A change process for prompts, templates, corpora and rules, with versioning and a clinical approval step. Without this, someone edits a template on a Friday to fix a complaint and nobody can reconstruct what changed.
- An on-call rota, and a decision about what happens when the system is down: does the process fall back to manual cleanly, or does the unit stop? Answer this before go-live, in writing, per unit.
Staffing that run-team costs real money, recurring, and it is usually not in the business case because the pilot did not need it. I would now put the run-team cost in the first business case even if it makes the payback look worse, because the alternative is a system that gets switched off and a board that will not fund the next one.
Integration debt comes due all at once
Pilots get permission to take shortcuts, and every shortcut is a loan. The nightly extract becomes a real-time interface. The read-only view becomes a write-back into the HIS, which is a different conversation with the HIS vendor entirely, with a different risk posture and usually a quote and a timeline. The single unit’s test instance becomes several production instances on different versions with locally modified masters.
The three that have cost me the most time:
- Write-back. Reading from the HIS is a project. Writing into it — an appointment, a note, an order — touches clinical record integrity, needs the HIS vendor’s involvement, and in practice sets your timeline rather than your own team’s capacity.
- The interface engine queue. There is usually one engine, one person who understands it, and a backlog of requests from finance, diagnostics and the insurance desk ahead of you. Your production date is their queue position.
- Identity. The pilot worked because somebody hand-reconciled the sample. Production needs the cross-reference layer to exist, which is a separate programme that should have started earlier.
There is a useful discipline here: at pilot design time, write down every shortcut taken and what the production version of it costs in weeks and rupees. Keep that list visible in the pilot review. It converts “the pilot succeeded” into “the pilot succeeded and here is the bill”, which is the conversation the board should be having.
Procurement is not a formality and it is not fast
Pilots run on a proof-of-concept letter, a small amount of discretionary spend, and goodwill. Production needs the full machinery: vendor empanelment, an information security review, a data processing agreement that survives your DPO’s reading, uptime and support commitments with teeth, an exit and data-return clause, and capital or operating approval at a level that goes to a committee meeting with a fixed calendar.
Assume that sequence takes a quarter and possibly two, and that the infosec review will surface something material — data leaving the country, a subprocessor nobody disclosed, logs retained indefinitely, a support model that requires production data access from outside. Each of those is resolvable. None is resolvable in a week, and all are cheaper to resolve before the clinicians have been promised a date.
The commercial shape itself often breaks the case. Per-conversation or per-seat pricing that is trivial in one unit becomes a significant recurring line across a group, and the cost curve goes the wrong way precisely as volume grows. Model the group-wide cost at the pilot stage. I have seen a genuinely good pilot die because nobody did, and the number that emerged at scale was not defensible against the alternative of three more people in the contact centre.
Change management, when the units do not report to you
In a multi-unit group, the corporate function proposes and the unit disposes. A unit head with an occupancy target and a monthly review does not care about your platform, and is right not to. Anything that adds a step at the front desk or in the OPD without an obvious local benefit will be absorbed for two weeks and abandoned in the third.
What moves a unit:
- The benefit is local and visible — fewer calls into the front desk, fewer no-shows in their OPD, a nursing shift that ends on time.
- The SOP is rewritten and the NABH documentation updated, so the new process is the documented process rather than an informal overlay on the old one. Until this happens, every staff rotation resets your adoption.
- Training reaches the actual operators — the front-desk staff on the evening shift, the OPD nurse, the contact-centre agents on the regional-language desk — not only the managers who attended the briefing.
- Somebody in the unit has it in their own objectives. Not a champion in spirit; a line in a review.
- The old path is closed, on a date, after the new one is demonstrably working. Parallel running forever is how good systems die quietly.
That last one needs care in a hospital. You cannot close a clinical fallback casually. But administrative fallbacks — the parallel register, the second booking sheet, the WhatsApp group that bypasses the system — have to be closed deliberately or the data will never be complete enough to trust.
The decisions to make before the pilot starts
This is the part I would do differently everywhere. Each of these is cheap to decide before and expensive to decide after.
- What production looks like. Which units, which channels, what volume, what the end state process is. If you cannot describe it, you are not piloting, you are experimenting — which is fine, but call it that and do not let anyone plan a rollout around it.
- Who the business owner is in production. A name and a function, agreed with them.
- What the run-team is and what it costs. In the first business case.
- The integration path and its dependencies. Specifically whether write-back is required, and the interface engine’s queue position.
- Group-wide unit economics. At full volume, not pilot volume.
- The kill criteria. What result, by what date, means you stop. Write it before you are emotionally invested.
- What the old process becomes. Retired, fallback, or permanent parallel — and who decides.
- Who signs off clinically, and what the review cadence is once it is live.
A pilot that cannot answer those eight is a demonstration. Demonstrations have their uses; just do not put them in the board pack as progress.
If you’re starting this next quarter
- Count your live pilots. If there are more than two or three, stop at least half. Pilot sprawl is the real problem in most groups, and it is consuming the scarce attention of your clinical champions.
- Pick the one with the clearest business owner — not the most impressive technology — and write the eight decisions down for it.
- Start procurement, infosec and the data processing agreement in parallel with the pilot, not after it. This alone saves a quarter.
- Get the run-team funded as part of the same approval. Present it as operating cost, not as project contingency, because that is what it is.
- Rewrite the SOP and the NABH documentation for the new process during the pilot, while the process is still being learned.
- Pick unit two on the basis of data readiness and a willing unit head, and treat it as the real test. Unit two is where you find out whether you built a system or a relationship.
- Set the date for closing the old administrative path, and hold it.
Pilots are easy because they ask nothing of the organisation. Production is the part where somebody has to put their name on it.
Questions people ask
Because a pilot tests the idea and production tests the organisation. The pilot runs in the friendliest unit, on a data extract, supported personally by someone who answers their phone at nine at night, with a vendor working at an attention level no contract will sustain. None of that survives contact with unit four. The gap is a set of commitments — an owner, a run-team, an integration path, a budget line — that a pilot is specifically designed to avoid making.
A named business owner who runs the process it sits inside — contact centre head for an enquiry agent, medical administration for documentation, the diagnostics lead for anything touching reports. The technology function is accountable to them on a defined service level, and a clinical authority owns the standard and the never-do list but not the operation. Three roles, one decision-maker, written down before the pilot. A steering committee owning it means nobody does.
A first line that can answer a nurse’s question at eleven at night, a second line with logs and a runbook for the five known failures, someone watching quality metrics weekly, a change process for prompts and templates with clinical approval, and an on-call rota with a written fallback per unit. It is recurring operating cost and it is rarely in the business case because the pilot did not need it. Put it in the first case even if the payback looks worse.
Every shortcut the pilot took, and each is a loan. The nightly extract becomes a real-time interface. The read-only view becomes write-back into the HIS, which involves the HIS vendor and their timeline. The hand-reconciled sample becomes a cross-reference identity layer that should have started earlier. Write down every shortcut at design time with what its production version costs in weeks and rupees. It turns “the pilot succeeded” into “the pilot succeeded and here is the bill”.
Reading from the HIS is a project. Writing into it — an appointment, a note, an order — touches clinical record integrity, needs the HIS vendor’s involvement, carries a different risk posture and usually arrives with a quote and a timeline you do not control. The interface engine adds a second constraint: one engine, one person who understands it, and a queue of finance and diagnostics requests ahead of you. Your production date is your queue position.
Assume a quarter and possibly two. Production needs vendor empanelment, an information security review, a data processing agreement that survives your DPO, uptime and support commitments with teeth, an exit and data-return clause, and approval at a committee with a fixed calendar. The infosec review will surface something material — data leaving the country, an undisclosed subprocessor, indefinite log retention. Start all of it in parallel with the pilot, not after.
Support commitments that hold at eleven at night in unit four, not the pilot team’s goodwill. A data processing agreement covering where data goes, which subprocessors touch it and how long logs are kept. Whether support requires production data access from outside the country. An exit clause with data return. And a commercial shape modelled at group-wide volume, because per-conversation pricing that is trivial in one unit can make the whole case indefensible at scale.
Because the cost curve goes the wrong way precisely as volume grows. A per-seat or per-conversation rate that looks like a rounding error in one unit becomes a significant recurring line across a group. I have seen a genuinely good pilot die because nobody modelled the group-wide number, and when it emerged it could not be defended against the alternative of three more people in the contact centre. Model full volume at the pilot stage.
The benefit has to be local and visible — fewer front-desk calls, fewer no-shows, a nursing shift that ends on time. The SOP and NABH documentation are rewritten so the new process is the documented one. Training reaches the evening-shift front desk and the regional-language agents, not only managers. Someone in the unit has it in their own objectives. And the old administrative path — the parallel register, the WhatsApp group — is closed on a date.
Eight. What production looks like — units, channels, volume, end-state process. Who the business owner is in production, by name. What the run-team is and costs. The integration path, including whether write-back is required and the interface engine’s queue position. Group-wide unit economics at full volume. The kill criteria. What the old process becomes. And who signs off clinically, with what review cadence. A pilot that cannot answer these is a demonstration.
A written statement of what result, by what date, means you stop. Write it before you are emotionally invested and before the clinicians have been promised a rollout. Without it, a pilot that met objectives nobody defined drifts into a permanent pilot, consuming your clinical champions’ attention. Kill criteria are also what a board should ask for, because they turn a status report into a decision the committee can hold someone to.
Two or three at most. Pilot sprawl is the real problem in most groups — a portfolio of successful pilots and nothing in production, each one drawing on the same scarce clinical champions and the same overloaded interface engine. Count your live pilots and stop at least half. Pick the one with the clearest business owner, not the most impressive technology, and write the eight pre-pilot decisions down for it.
Because unit one had the best IT, the most cooperative unit head and your personal attention. Unit two has a different HIS version, locally modified masters, a unit head with an occupancy target who does not care about your platform, and no one answering their phone at nine at night. Pick it on data readiness and a willing unit head, and treat it as the test. Unit two is where you learn whether you built a system or a relationship.
The clinical standard and the never-do list, and the approval step in the change process — so that when someone edits a template on a Friday to fix a complaint, a clinician has signed it and the change is versioned. They do not own uptime or the operation. Clinical sign-off and a review cadence once the system is live should be decided before the pilot, not negotiated after clinicians have been promised a date.
