Improvement and impact are different claims
A healthcare intervention may be associated with better access, quality, experience, equity, timeliness, utilization, or cost. An observed improvement is important, but it does not by itself show what caused the change.
The causal question is what would have happened to the same eligible population, during the same period, if the intervention had not been introduced. Because that alternative cannot be observed directly, the evaluation must construct a credible comparison.
CAUSAL QUESTION
What changed because of the intervention, relative to what would otherwise have happened?
Define one decision first
Start with the action the evidence must inform. The decision may be to continue a program, expand it, change eligibility, modify delivery, redirect resources, or stop.
A broad question such as “Did the program work?” is not yet evaluable. Specify:
- the intervention and version;
- the eligible population;
- how exposure is determined;
- the primary outcome;
- the observation window;
- the unit of analysis;
- the comparison strategy; and
- the decision threshold.
These elements should be recorded in a Measurement Contract before the result is interpreted.
Draw the causal structure
A causal diagram makes assumptions visible. It distinguishes variables that influence both exposure and outcome from variables caused by the intervention. This matters because adjusting for the wrong variable can introduce bias rather than remove it.
The diagram should reflect operational knowledge from people who understand eligibility, referral, access, delivery, and outcome recording. Statistical convenience is not a substitute for that knowledge.
Choose a comparison that fits the rollout
Randomization provides a strong basis for attribution when it is feasible and ethical. Many real interventions require quasi-experimental designs.
Useful designs can include:
- a phased or randomized rollout;
- difference in differences with examined pre-trends;
- matching or weighting with adequate overlap;
- an interrupted time series with sufficient history;
- a threshold or discontinuity created by eligibility rules; or
- a justified natural experiment.
No method repairs a comparison that does not represent the relevant counterfactual. Diagnostics should test the assumptions on which the design depends.
Measure implementation without confusing it with effect
Adoption, reach, fidelity, timeliness, and completion explain whether the intervention operated as intended. They may also explain why effects vary. They are not automatically the outcome.
Capture intervention version and exposure timing. If delivery varies, record enough information to distinguish assignment, actual exposure, and implementation fidelity. Otherwise, the analysis may compare labels instead of meaningful treatment differences.
Predefine refusal conditions
An evaluation should state when it cannot support the intended claim. Examples include inadequate sample size, missing required context, weak overlap, failed balance, unstable measurement, concurrent changes that cannot be separated, or a comparison period too short to establish trend.
Refusal is part of the method. It prevents an operational dashboard or exploratory association from being relabeled as causal evidence.
Report the result at the supported level
The result may support one of four levels:
- Operational: what happened in the observed data.
- Comparative: how defined groups or periods differed.
- Causal: the estimated effect under a defensible design.
- Withheld: the planned claim is not supported.
The result should include uncertainty, sensitivity, subgroup findings, limitations, and data lineage. When cost is part of the decision, calculate causal ROI rather than modeled savings.
Make the evidence reusable
An Impact Passport keeps the claim attached to the intervention version, Measurement Contract, data window, analysis, diagnostics, value inputs, and approvals. This lets a program owner, finance leader, or reviewer understand not only the result, but also what the result is allowed to mean.
Method context
The CDC Program Evaluation Framework distinguishes impact evaluation, which compares outcomes with and without a program, from other useful evaluation questions. It also treats collaboration, fair evaluation, rigor, independence, transparency, and use of findings as parts of the evaluation process.