AI Governance

Operational AI evidence belongs alongside documentation

Why healthcare AI needs scoped operational evidence alongside policies, and what a signed record can and cannot establish about a covered event.

7 min read
Joe Braidwood
Joe Braidwood
Co-founder & CEO
December 2025 · 7 min read

The practical shift: Policies, procedures, and assessments still matter. Consequential AI also benefits from signed, scoped records connecting an intended rule to what a configured control reported. Those records add integrity and attribution; they do not independently prove safety, effectiveness, complete coverage, or compliance.

Documentation vs. Evidence

The two answer different questions, and the gap between them opens up the moment somebody asks about a single request rather than about the program as a whole:

Documentation Says

  • “We have guardrails”
  • “We monitor for bias”
  • “We log all requests”
  • “We have human oversight”

Evidence Can Support

  • “Here is the signed record guardrail X reported”
  • “Here’s the bias test result from timestamp Y”
  • “Here’s a verifiable record of request Z”
  • “Here is the signed review event recorded at time T”

Documentation records intent. Operational evidence records selected events and claims. The signature can make later alteration detectable, while source truth, effectiveness, and coverage still need separate evidence.

Why AI changes the equation

Traditional software tends to give teams reproducible behavior under the same code and inputs. AI systems vary more, fail in ways that are harder to see from the outside, and depend heavily on data, prompts, and model versioning.

Four properties drive that difference:

  • Non-deterministic outputs: the same input can produce different outputs
  • Emergent behaviors: models exhibit capabilities (and failures) not explicitly programmed
  • Continuous drift: behavior changes over time, sometimes subtly
  • Context sensitivity: outputs depend on complex combinations of inputs

With AI, design documentation alone does not establish operational behavior. Reviewers need scoped records of what the system reported, plus testing and coverage evidence appropriate to the decision.

The four pillars of AI evidence

Based on the questions that show up most often in regulation, procurement, and incident review, we think four capabilities matter most:

1. Guardrail Execution Trace

Signed traces showing what configured controls reported, in what recorded sequence, and with which result. The trace supports integrity for covered fields; it does not by itself prove event truth, control effectiveness, or complete routing.

2. Decision Rationale

Selected context references, commitments, configuration identifiers, and declared scope tied to an output. A data-minimizing record is not a complete explanation of model reasoning or a full forensic reconstruction.

3. Independent Verifiability

Cryptographically signed, tamper-evident records whose supported fields a third party can check. Organizational independence and key custody must be established outside the artifact.

4. Framework Anchoring

A declared mapping between selected record fields and relevant objectives in ISO 42001, NIST AI RMF, or applicable EU AI Act provisions. A mapping supports review; it does not establish that a requirement is satisfied.

The boundary: These pillars do not replace documentation. They add independently checkable records for configured, in-scope actions, with coverage and control effectiveness evaluated separately.

What this looks like in practice

For a healthcare AI system processing clinical notes, evidence-grade operations would produce:

  • Scoped request record: a signed account of selected processing steps on the configured path
  • Redaction-control result: a record of what the configured redaction control reported, with effectiveness tested separately
  • Model-version commitment: a digest bound to the declared model-version field, not independent proof of the deployed model
  • Guardrail event records: results for covered controls, accompanied by routing and denominator evidence
  • Review timeline: a sequence of covered records and explicit gaps, not a claim of complete chain of custody

For high-stakes AI deployments, this is the kind of operational evidence buyers, auditors, and regulators increasingly ask for when something goes wrong.

The regulatory convergence

Several frameworks push in the same direction, even if they use different language:

  • EU AI Act Article 12 requires automatic recording of events for covered high-risk systems
  • Colorado’s automated-decision law (SB 26-189), which repealed and replaced the 2024 Colorado AI Act before it took effect, centers on documentation, pre-use notice, and disclosure for covered automated decision-making technology (ADMT), with substantive duties commencing January 1, 2027
  • NIST AI RMF structures governance around mapping, measuring, managing, and governing risk
  • ISO 42001 is a management-system standard rather than a product-safety certificate

The common thread is a push toward operational evidence rather than written policy alone.

The competitive advantage

In practice, organizations that build evidence infrastructure early are better positioned for:

  • Faster security reviews: evidence is more compelling than documentation
  • Incident response: there are records to review when something goes wrong
  • Regulatory readiness: records are easier to connect to the relevant control set
  • Internal governance: oversight decisions can be tied back to operating evidence

Teams still relying on documentation alone are likely to have a harder time in reviews, diligence, and incident response because they cannot easily connect policy claims to operating records.

The path forward

Moving from documentation to evidence requires infrastructure changes:

  • Configured-path records: capture in-scope decisions and state the denominator and exclusions
  • Cryptographic attestation: sign covered fields so later alteration can be detected
  • Independent verification: enable third parties to check supported cryptographic properties without asking the operator to declare a pass; source truth and trust still require separate evidence
  • Framework mapping: connect evidence to specific regulatory requirements

None of this is a compliance checkbox. For healthcare and other high-stakes uses, relying on policy documents alone is increasingly hard to defend.

For the complete technical framework, read our white paper.

Primary sources

Pango waving

The complete framework

Our white paper “The Proof Gap in Healthcare AI” provides the full technical analysis of evidence infrastructure, including architecture patterns and vendor assessment checklists.

Read the White Paper