Field notes
From AI governance in healthcare to defensible evidence.
Governance documents say what should happen. After an adverse event, somebody asks what did happen, which safeguards ran, and who outside the system can confirm it.
A lawyer, a doctor and an insurer walk into a room. They all want to know the same thing, and each of them asks for it differently. The lawyer asks, “what can we prove?” The doctor asks whether the error was theirs, or the tool’s. The insurer asks which line of coverage this falls under.
What does a defensible audit trail actually look like? How do you attribute liability when AI is involved?
I have had a dizzying number of conversations over the past few weeks with risk and compliance leaders, medical malpractice attorneys, insurers, and the CIOs and CMIOs of healthcare organizations trying to answer those two questions. At the CHAI Legal Summit in Boston in September, I led a roundtable on AI insurability and defensibility with attorneys and health system leaders. Here is what I took away.
One distinction kept coming up, and it is where AI governance in healthcare now has to go: monitoring is not the same as evidence. Monitoring can tell you that a system is behaving differently, that an error rate is rising, or that a control is firing. Evidence needs to let you establish what happened in a particular event, which safeguards actually operated, who intervened, and whether that record can be independently trusted.
Monitoring
Tells you that a system is behaving differently.
- An error rate is rising
- Outputs are drifting from the baseline
- A control is firing more often, or less, than expected
Evidence
Establishes what happened in one particular event.
- Which safeguards actually operated
- Who intervened, and what they decided
- Whether the record can be independently trusted
I · The unit of analysisBeyond the model and the benchmark
For a long time, AI governance in healthcare has focused on model accuracy and performance against benchmarks. It is becoming clear that this matters and that it is the tip of the iceberg. Defensibility is shifting from “was it evaluated and approved” to “can you show how it was governed, validated, and monitored after deployment.”
The unit of analysis is changing. Model performance still matters, but organizations increasingly have to understand whether the system around the model worked: the workflow, controls, escalation paths, human review, permissions, and response to failure. And when something does go wrong, governance documents are not enough. Someone will eventually ask for evidence that those controls actually operated.
It was refreshing to see the conversation move to operational and system accuracy: what the controls are, who the humans reviewing it are, what happens when something goes wrong, and what you would want to be able to prove after the fact.
Clinical AI tools increasingly run on multiple models or agents, wrapped in workflows, guardrails and human-in-the-loop steps. A model can perform well on a benchmark and the use case can still fail and lead to a bad outcome, because a safety control didn’t work as designed. A model can correctly flag a patient message as high risk, but if the escalation rule that routes it to a clinician doesn’t fire, the patient still falls through.
One high-risk patient message
- 01 A patient message arrives in the inbox Input
- 02 The model flags it as high risk Model · correct
- 03 The escalation rule routes it to a clinician Control · did not fire
- 04 Nobody sees the message Outcome · patient falls through
The model was accurate. The system failed.
II · After an adverse eventThe post-loss record
After an adverse event, the questions that matter are these.
- What did the AI system show at the moment of the decision?
- What did it recommend, and what did it actually do?
- Was it an error on the AI vendor’s side, or the deployer’s side?
- Which named clinician accepted it, overrode it, or reopened it?
- Does the organization own that record, or does it have to go ask the vendor?
None of those are accuracy questions.
The field is already moving in this direction. Guidance from the Joint Commission and CHAI recommends that organizations develop systems to track the performance, outcomes and adverse events from using AI tools, and advocates for voluntary, blinded reporting of AI safety events to independent organizations like patient safety organizations.1 But reporting an event and reconstructing it are two different things.
Other questions remain open: how to establish what actually happened from an audit log (an audio or recording failure, say, as against a model or dataset issue), where the line for clinical decision-making sits, how consent and privacy are handled, and what the training data was.
III · Pricing the riskWhat insurers are underwriting
Insurers don’t want to guess at how to price AI risk, and right now they have very little to go on. The quality of the evidence an organization can produce is increasingly relevant to how insurers understand, and eventually price, that risk.
Underwriters want to understand how a model was built, how a representative dataset was assembled (or why a smaller one was appropriate), and what makes a provider low risk, including the type and quality of training clinicians have had on using the AI tool. They also want to see governance built on a strong foundation, ongoing oversight, and how an organization’s approach has changed over time. Showing your risk mitigation, the rationale behind it, and how it has evolved can help with pricing.
The core underwriting question is simple. If the technology doesn’t work as intended, what does it cost, and what is being promised?
IV · CoverageInsurance coverage is evolving
Medical malpractice, technology E&O, cyber, and other liability policies can respond to different parts of an AI-related loss, depending on the facts and the policy language. Bundled products exist, and they may reduce some of the gaps and finger-pointing, but an AI incident can still implicate multiple policies, insureds, and theories of liability.
Policies that can respond
- Medical malpractice
- Technology E&O
- Cyber
- Other liability
- Developer
- Health system
- Clinician
Parties who may be on the hook
There is no settled answer on who controls the defense, who partners, or how liability gets divided among the developer, the health system and the clinician in a way that reduces finger-pointing.
V · StandardsAI governance in healthcare has no agreed standard yet
There was discussion on whether healthcare AI will get its own SOC 2 or HITRUST equivalent. Several people said they would love to be told what to do and they will do it, but it isn’t one size fits all.
Part of the problem is that we lack an agreed-upon standard for what evaluating an AI system should look like. PACT AI recently published a helpful piece, An AI Evaluator Is Not an Auditor, that separates four kinds of evaluation we tend to lump together.8
Investigation
Reconstructs what happened after an incident.
Evaluation
Measures what a system does under test conditions. Not only model accuracy, but operational accuracy.
Assessment
Judges whether practices are adequate for a purpose.
Audit
Reaches a conclusion against criteria fixed in advance.
An audit can’t conclude anything until the criteria exist.
So what standard of care do we measure against when none exists yet? Some have started proposing licensure-style frameworks for clinical AI,34 but those are just proposals.
There are signs of progress. In June, the Joint Commission launched its Responsible Use of AI in Healthcare (RUAIH) certification, with standards organized around governance; data management; risk and bias reduction; monitoring, evaluating and validating safety performance; and transparency, education and training.2
Certifications like RUAIH and ISO 42001 can establish that an organization has a responsible AI program, governance processes, safeguards and monitoring in place. That is still different from establishing what happened during a particular AI-mediated event, or whether a specific control operated as intended. We need both.
Utah’s Office of Artificial Intelligence Policy recently published Evidentiary Expectations for Healthcare AI Sandbox, which could serve as a guidepost from a regulatory standpoint.5 Section 5.3.3 calls for every applicant to maintain process logging robust enough to support internal and external audits after a severe incident, including traceability and retention of chat logs and reasoning traces to support harm attribution. The guidance also calls for auditable documentation of model changes and notification of severe incidents within 24 hours, and states a preference for third-party validation of quality assurance data. The EU AI Act’s record-keeping requirements for high-risk systems point the same way.6
That matters because it begins to move from “show us what you tested” toward “show us what happened in operation.”
This shift gives more definition. Health systems still need a working baseline for what “good enough” looks like in practice: which controls are required at minimum, what a baseline set of operational controls looks like for each use case, what gets recorded, and who reviews it. And the evidence has to be captured now, continuously, so it is there when the criteria catch up.
VI · After deploymentFrom monitoring to evidence
Traditional performance monitoring tells you how a system is performing over time. It doesn’t necessarily tell you what happened in a specific encounter: which controls ran, what they caught, whether a clinician reviewed the output, and what was done about it. It also doesn’t give a third party a way to check that the record hasn’t been altered.
A control existing is not the same as knowing that it operated as intended. And post-deployment conditions do not remain static.
Kopanitsa found that validation-era performance did not remain stable after deployment, even when the models themselves had not been intentionally changed. Calibration drift, changes in data availability, latency and evolving clinical workflows all contributed to divergence from the original baseline, with some operational signals providing earlier warning than outcome measures.7
Post-deployment assurance cannot stop at “did the model change?” Organizations also need to know whether the system’s real-world behavior, inputs, workflow and use remain aligned with the baseline intent and the conditions under which it was approved.
A model’s account of its own behavior is not independent evidence of what happened. Logs can be evidence, but their evidentiary value depends on four things.
- Origin
- What generated it
- Coverage
- What it captures
- Integrity
- Whether it can be altered
- Independence
- Whether anyone outside the system can verify it
VII · AheadWhat’s next
-
OVERT, the open evidence standard we developed, has been placed under the stewardship of CHAI and the AIGovOps Foundation, where the community and domain experts can continue to shape it. Read why we gave up our vote, or read the announcement.
-
We will also be starting the Lloyd’s Lab accelerator program, to explore the role of runtime evidence in measuring and pricing AI risk.
There are still many questions to answer. What evidence standard should health systems, vendors, insurers and regulators share? How should runtime evidence support audits, accreditation, licensure, regulatory review and insurance? How do governance programs scale without slowing deployment, or letting review quality slip?
One thing is becoming clearer. The next phase of responsible AI will not be defined only by whether organizations have governance programs, whether models perform well in testing, or whether systems are monitored after deployment. It will also be defined by whether we can establish what actually happened when AI was involved.
When a clinical AI system contributes to a bad patient outcome, how should fault be allocated among the treating clinician, the deploying health system and the developer? What might a safe harbor look like? What was the system authorized to do, and what did it actually do? Which safeguards ran? Can someone other than the system itself verify the answer?
Monitoring helps us see AI. Evidence is what will let us defend it.
References
- Joint Commission and Coalition for Health AI (CHAI). Guidance on Responsible Use of AI in Healthcare. 2025. Addresses ongoing local quality monitoring, validation and testing, adverse-event reporting, and risk-based post-deployment monitoring. Joint Commission
- Joint Commission. Responsible Use of AI in Healthcare (RUAIH) Certification. Launched June 1, 2026. Joint Commission
- Bressman E, Shachar C, Stern AD, Mehrotra A. Software as a Medical Practitioner — Is It Time to License Artificial Intelligence? JAMA Internal Medicine. 2026;186(1):5–6. doi:10.1001/jamainternmed.2025.6132. PubMed
- Bergman A, Wachter RM, Emanuel EJ. A Licensure Framework for Autonomous Clinical AI. JAMA. 2026;335(20):1751–1754. doi:10.1001/jama.2026.5483. JAMA Network
- Utah Office of Artificial Intelligence Policy. Evidentiary Expectations for Healthcare AI Sandbox. 2026. Describes continuous post-deployment quality assurance and empirical evidence collection as central to the sandbox approach. Utah Department of Commerce
- European Union. Regulation (EU) 2024/1689, Article 12, Record-keeping. High-risk AI systems must support automatic event logging to provide traceability, facilitate post-market monitoring, and monitor operation. EUR-Lex
- Kopanitsa G. Validation is not enough: longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. 2026;5(7):e0001534. doi:10.1371/journal.pdig.0001534. PLOS
- PACT AI. An AI Evaluator Is Not an Auditor. The source for the distinction among investigation, evaluation, assessment and audit. PACT AI