Pressure gauge fitted among metal pipes and valves on a machine

Measuring an AI project when nothing clinical is being claimed

16 min read

Measuring AI projects that make no clinical claim means proving an operational or commercial change, not a health outcome. Set a baseline before launch, compare against a group or period without the tool, sample output quality by hand, track adoption and full running cost, and agree in advance what result would make you stop. Keep every measure away from implying clinical benefit.

The easiest AI project to approve in a hospital is one that stays away from clinical care. A model that summarises call notes, drafts content, sorts enquiries or reconciles TPA paperwork does not need an ethics committee, and nobody worries it will harm a patient directly. So it gets approved on a slide, launched on enthusiasm and then judged on anecdote.

Six months later someone in finance asks whether it worked. The team shows usage charts and a few happy quotes from users. Nobody can say what changed, compared with what, at what cost. The project either limps on because stopping it feels awkward, or it is cut in a budget round without anyone knowing whether it was good.

Measuring AI projects that make no clinical claim is its own discipline. The outcomes are operational and commercial, the comparisons are messy, and there is a constant temptation to borrow clinical language to make the result sound bigger. This is how I would approach it.

Why non-clinical does not mean unmeasurable

Clinical AI has an established evaluation culture, with validation studies, sensitivity and specificity, and regulatory expectations. Non-clinical AI has almost nothing equivalent in most hospitals. Teams fall back on vendor dashboards, which measure the tool’s activity rather than the hospital’s benefit.

But operational work is often easier to measure than clinical work, not harder. The outcomes are closer in time and easier to observe. How long enquiries wait. How many claims are returned for missing documents. How many hours a coordinator spends on a report. How many draft articles need major rework. These are countable, and most hospitals already count some of them.

What makes measurement hard is not the outcome. It is the discipline around it: deciding what to measure before launch, collecting a baseline, and resisting the urge to change the measure once results come in. That is a management problem, and it is the same one I describe in why hospital AI pilots never reach production.

Start with the decision, not the dashboard

Before choosing any measure, write down the decision the measurement will inform. Usually it is one of three: expand the tool to more units or teams, keep it running as is, or stop. Sometimes it is a fourth, change the design and test again.

Then ask who makes that decision and what they need to see. The CFO wants the net effect on cost or revenue with a credible comparison. The unit head wants to know whether the team’s workload and service actually improved. The operations lead wants to know whether quality held up. Each needs a different measure, but all of them should be agreed before launch.

Writing the decision down first has a useful side effect. It exposes projects that have no real decision attached, where the tool was bought because a peer hospital had one or a vendor offered a free pilot. Those projects are not wrong in themselves, but they need a sponsor who will own the result before any measurement is worth doing. Without one, the evaluation will be read by nobody.

I also write down the result that would make us stop. This is the part teams resist most. But if there is no pre-agreed failure threshold, every result becomes “promising” and the project never ends. A stop rule written in advance is kinder to everyone, including the team who built the thing.

The four families of measures that matter

For most non-clinical AI work in hospitals, the useful measures fall into four families. I try to have at least one from each.

  • Outcome: the operational or commercial change the project exists for, such as time to first response, claim rejection for documentation reasons, turnaround on discharge paperwork, or conversion from enquiry to appointment.
  • Quality: whether the output is good enough, measured by sampling and human review, not by the model’s own confidence.
  • Adoption: whether the people meant to use it actually do, and whether they keep using it after the novelty fades.
  • Cost to run: the full cost, including licences, integration, review time, supervision and the internal team maintaining it.

A word on commercial outcomes. When a project is meant to lift revenue, such as a tool that helps the contact centre convert more enquiries, it is tempting to credit every extra appointment to the tool. Revenue moves for many reasons: a new doctor joins, a competitor closes a department, a campaign runs. Measure the step the tool actually touches, such as the conversion rate of enquiries it handled against those it did not, and let finance translate that into revenue with their own assumptions. The claim is smaller, and it survives scrutiny.

Outcome without quality is dangerous: a faster process producing worse work. Quality without adoption is academic. And outcome without cost to run is how projects look profitable in a pilot and expensive at scale.

Baselines and comparisons for measuring AI projects

A result only means something against a comparison. The most common mistake I see in measuring AI projects is comparing the post-launch period with nothing at all, or with a vague memory of how things used to be.

Collect a baseline before launch

Measure the outcome for several weeks before the tool goes live, using the same definitions you will use afterwards. If the process is not currently measured, that is itself a finding, and you may need a short manual sampling period to establish where you are starting from.

Find a fair comparison

Before and after comparisons are weak on their own, because hospitals have seasons. Enquiry volumes, admission patterns and staffing change across the year, and a tool launched in a quiet month will look brilliant. Where possible, compare against a group that did not get the tool at the same time: another unit, another shift, another team, or a random share of the work that stays on the old process. Even an imperfect comparison group is far better than none.

Hold the definitions still

Decide exactly how each measure is calculated, write it down, and do not change it during the evaluation. Redefining “response time” halfway through, even for good reasons, makes the before and after incomparable. This is ordinary measurement hygiene, the same discipline that applies to attribution in healthcare more broadly.

Quality needs human sampling

Language model output can look excellent and be wrong. So quality has to be measured by people reading a sample of real output against a clear standard, on a regular rhythm.

For each project, define what “good enough” means in plain terms. A call summary is good if it captures the reason for the call, the action agreed and the follow-up owner, with nothing invented. A drafted article is good if an editor can publish it after light changes and nothing factual needs correcting. A document check is good if it flags the same missing items a trained coordinator would.

Then have a small group review a random sample each week, score it against that standard and note the types of error. Random matters. Reviewing only the cases people complain about tells you about complaints, not quality. And watch the trend: quality often drops after launch, when the tool meets inputs that were not in the test set, or when the vendor updates the underlying model.

Who does the sampling

The reviewers should be people who know what good looks like for the task, usually experienced staff from the team that does the work, plus someone from outside that team for independence. Give them a short scoring sheet, a fixed sample size each week and protected time to do it. If sampling is squeezed into spare moments, it stops within a month, and the quality measure silently disappears from the evaluation. Rotate reviewers occasionally so that one person’s standards do not become the only standard.

Adoption and the hidden cost of review

A tool that people avoid has no effect, however good it is. Track how many of the intended users use it, how often, and whether use holds up after the first month. Falling adoption is an early signal that something is wrong: the tool is slow, the output needs too much fixing, or it does not fit how the work is actually done.

The flip side is review load. Most responsible designs keep a human checking the AI output, and that checking takes time. If an assistant saves a coordinator time drafting a summary but the supervisor now spends that time reviewing summaries, the net saving may be small. Measure the time spent reviewing, correcting and supervising, not just the time the tool appears to save.

This is also where early numbers mislead. In the first weeks, enthusiastic users drive adoption and the review load is light because volume is low. At scale, both change. The pilot that has to work before automation scales is one that has been measured under realistic volume, not only in a friendly corner of the hospital.

Keeping clinical language out of the results

Non-clinical projects have a habit of acquiring clinical claims as they are presented upward. A tool that speeds up discharge paperwork becomes “improving patient flow”, then “reducing length of stay”, then, in someone’s board slide, “better patient outcomes”. None of those were measured, and some of them would need clinical evidence to claim.

I hold a simple rule: report only what was measured, in the words it was measured in. If the project reduced the time discharge summaries wait for typing, say that. If you believe it also affected bed turnaround, that is a separate hypothesis requiring its own measurement, and it should be described as such. Never imply a health benefit from a project designed and evaluated as an operational one.

The same caution applies to patient-facing tools. A chatbot that answers logistics questions can be measured on resolution and satisfaction. It should not be described as improving access to care unless someone has actually measured access. I learnt this the uncomfortable way with containment versus resolution, where a flattering number concealed a worse experience.

Reporting upward without overclaiming

When the evaluation period ends, the report should be short and specific. State the decision it informs. Show the outcome against the baseline and the comparison group, with the definitions. Show quality from sampling, adoption over time and the full cost to run. Then give a recommendation: expand, continue, redesign or stop.

The CFO will ask about the counterfactual, and they are right to. As I argue in what a digital head owes the CFO, the digital leader’s credibility depends on showing results that survive a sceptical reading. It is better to report a modest, well-evidenced result than a large, fragile one.

Measurement also does not end when the evaluation does. A tool that passed its evaluation still needs a lighter, permanent version of the same checks: a monthly quality sample, adoption tracking and an outcome measure on the operational dashboard. Models change, processes change, and staff change. The monitoring that continues after the decision is what catches the slow decline that otherwise shows up only as a complaint.

And report the failures. A project stopped on a pre-agreed rule is a sign that the organisation evaluates AI properly. Boards generally trust leaders who can show what they stopped as well as what they scaled.

The measurement plan to write before kick-off

Before the next AI project starts, write a one-page measurement plan and get it signed by the project sponsor and finance. It should name the decision the evaluation will inform and who makes it, the outcome measure with its exact definition, the baseline period and how the baseline will be collected, and the comparison group, with honest notes on its weaknesses.

Add the quality standard and the sampling routine, including who reviews and how often. Add the adoption measures and the full cost to run, with review and supervision time included. State the evaluation period. And write the stop rule in plain words: the result below which the project ends or is redesigned.

Then collect the baseline before anyone switches the tool on. That single step is the one most often skipped, and it is the one that decides whether you will ever be able to say, with a straight face, that the project worked.

Questions people ask

What does measuring AI projects mean when nothing clinical is claimed?

It means evaluating AI tools used in operations, marketing, contact centres or administration on operational and commercial outcomes rather than health outcomes. The approach sets a baseline before launch, compares results with a group or period without the tool, samples output quality by hand, tracks adoption and the full cost to run, and agrees in advance on a result that would lead to stopping or redesigning the project.

As CFO, what should I insist on before approving a project?

Ask for a one-page measurement plan before launch. It should name the decision the evaluation will inform, define the outcome measure, describe how the baseline will be collected, identify a comparison group and state a stop rule. Also ask for the full cost to run at scale, including review and supervision time. If the team cannot write this page, the project is not ready.

How long should an evaluation period last?

Long enough to cover normal variation in the work and to get past the early enthusiasm of the first weeks. For most operational tools, that means a few months rather than a few weeks, with a baseline of similar length before launch. Seasonal patterns matter in hospitals, so an evaluation that covers only a quiet period can make a tool look better than it is.

Why not just use the vendor’s dashboard?

Vendor dashboards usually measure the tool’s activity: messages handled, documents processed, users logged in. They rarely measure whether the hospital’s process improved, whether quality held, or what it cost in review time. They are useful for monitoring, but the evaluation should be based on your own definitions and data, agreed before launch, so that the result answers your question rather than the vendor’s.

What is a comparison group and do we really need one?

A comparison group is a unit, team, shift or share of work that continues without the tool during the same period. It shows what would have happened anyway, separating the tool’s effect from seasonal swings or other changes. It is not always possible to have a perfect one, but even an imperfect comparison is far more convincing than a simple before and after.

How do we measure output quality?

Define in plain terms what good output looks like for the task, then have people review a random sample of real output against that standard each week, recording scores and types of error. Random sampling matters, because reviewing only complaints tells you about complaints. Watch the trend over time, since quality can drop when the tool meets new kinds of input or the underlying model is updated.

As a unit head, what matters most to me?

Whether your team’s workload and service actually improved, and whether quality held up. Ask to see adoption among your staff over time, the time spent reviewing and correcting AI output, and the outcome measure for your unit compared with the baseline. If your team quietly stopped using the tool, that tells you more than any headline figure in the report.

As medical director, why should I care about a non-clinical project?

Because non-clinical projects can drift into clinical claims as they are presented upward, and because some operational tools touch clinical workflows indirectly, such as discharge paperwork or scheduling. Your role is to make sure results are reported in the terms they were measured, without implying health benefits, and to flag any point where an operational tool is starting to influence clinical decisions.

What is a stop rule?

A stop rule is a result, agreed before launch, below which the project will end or be redesigned. It might be an outcome that does not improve against the comparison group, quality below the agreed standard, or adoption that falls away. Writing it in advance prevents every result from being described as promising and lets the team stop a project without it feeling like a personal failure.

How do we account for the time spent reviewing AI output?

Measure it directly. Ask reviewers and supervisors to log time spent checking, correcting and escalating AI output for a sample period, or estimate it from a time study. Include it in the cost to run. Many projects save time in one role and add it in another, and the net effect is only visible when both sides are counted.

What should the board see?

A short report showing the decision, the outcome against baseline and comparison, quality from sampling, adoption over time, full cost to run and a recommendation. Boards also benefit from seeing which projects were stopped under pre-agreed rules, because it shows that AI spending is being evaluated properly rather than justified after the fact. Keep it to a page or two, and resist adding usage charts that do not bear on the decision.

What does IT need to provide for measurement?

IT usually needs to make the underlying data available: timestamps, queue records, document status, usage logs and the ability to separate work handled with and without the tool. They may also need to support random sampling of outputs for review. Agree these data needs before launch, because retrofitting logging after go-live is slow and the baseline may already be lost.

Can we claim patient experience improvements?

Only if you measured patient experience directly, for example through surveys or complaint data for the relevant process, with a baseline and comparison. Faster response times may well improve experience, but that is a hypothesis until measured. Report what you measured in the terms you measured it, and describe wider benefits as expected or possible rather than proven.

What if the results are mixed?

Mixed results are common and useful. Perhaps the outcome improved but review load rose, or one unit benefited and another did not. The report should say so plainly and recommend a specific next step, such as redesigning the review process or expanding only where it worked. Mixed results reported honestly build more credibility than a tidy success story that later unravels.

Free download

Get the Hospital Digital Growth Audit

A 25-point self-assessment across AI operations, growth & CRM, launches, leadership, and PR. Confirm your email and it arrives in your inbox, along with the full Tools & Checklists set. Occasional notes after; unsubscribe anytime.