Evaluate the deployment, not the category

“AI” is not a stable intervention. A deployment includes a product version, model, configuration, eligible users, workflow position, adoption pattern, human review process, and operating period. If any of those change, the treatment may change.

An ROI evaluation should identify the exact deployment that produced the evidence. General claims about the model or product category are not substitutes for results in the buyer’s environment.

START HERE

What changed for whom when this version entered this workflow?

Separate technical performance from organizational impact

Technical metrics may include accuracy, latency, retrieval quality, task completion, escalation, and error rates. Adoption metrics may include active users, sessions, acceptance, overrides, or calls handled.

These measures help determine whether the system operated. Organizational outcomes address whether the deployment improved something the buyer values, such as access, timeliness, quality, experience, workload, equity, utilization, or cost.

The evaluation needs both. A product cannot create the intended effect if it is not used or does not perform, but good technical performance does not guarantee an organizational outcome.

Record exposure before the outcome

The intervention record should include product and model version, eligibility, assignment, actual exposure, timing, intensity, exceptions, and meaningful workflow changes. Important contextual factors should be recorded before assistance can affect them.

This protects the time order required for causal reasoning and reduces the risk that post-intervention information is treated as a baseline characteristic.

Use rollout variation deliberately

A deployment often creates natural comparison opportunities. Sites, teams, or users may start at different times. Eligibility may be phased. Capacity may constrain access. Some groups may remain on the prior workflow temporarily.

When documented and ethically managed, that variation may support randomized rollout, difference in differences, matching, weighting, or interrupted time series. The evaluation should plan for comparison quality before universal adoption removes the opportunity.

Test for hidden implementation costs

AI costs extend beyond subscription or usage fees. Include integration, data preparation, workflow redesign, training, human review, exception handling, security, privacy, quality monitoring, model updates, and vendor management.

If AI saves time in one role but creates review or correction work elsewhere, both changes belong in the cost and outcome record.

Demand proportional claims

Common unsupported leaps include:

  • activity to outcome;
  • user satisfaction to clinical or operational effect;
  • reported time saved to cash savings;
  • before-and-after change to causal attribution; and
  • one customer’s result to every future deployment.

A pre-approved Measurement Contract defines what evidence is required for each claim level. Failed diagnostics, missing exposure, inadequate comparison, or unstable implementation should limit or withhold the claim.

Make vendor and buyer inputs inspectable

The vendor may be best positioned to provide version, usage, technical performance, and implementation records. The buyer controls organizational outcomes, operational context, approved cost inputs, and the decision the evaluation must inform.

Neither side should be able to silently redefine the outcome or value after seeing results. The final Impact Passport should identify data ownership, analysis responsibility, approvals, limitations, and permitted external use.

Evaluate before renewal, and design before expansion

An existing deployment may be evaluable when baseline, rollout, exposure, comparison, outcome, and cost evidence exist. If they do not, the most useful output may be a prospective measurement plan for the next contract period or expansion phase.

The goal is not to guarantee a positive ROI. It is to make the renewal or expansion decision more credible.

Method context

The NIST AI Risk Management Framework treats measurement and evaluation as continuing parts of AI risk management. The NIST AI Resource Center also provides testing, evaluation, verification, and validation resources. These assurance questions complement causal outcome and ROI evaluation. They do not replace it.

Verify one healthcare AI deployment →