GLACIS
Platform
Solutions
AI in production Supervise the AI you run, with a signed record others can check AI you sell Show customers how your AI is controlled, in a record they can check AI you oversee Read the signed record from a vendor’s AI instead of taking their word for it All solutions Workflows and industries
Standards
Operational evidence How signed records become evidence a reviewer can check Sample record Every action, from its rule to a record anyone can check OVERT standard The open record format, published June 2026 Verify a record Check a signed operational record yourself, in your browser
Resources Company Talk to us Start free
GLACIS

Navigate

Home PlatformRules, control decisions, and records anyone can verify Resources Company

Solutions

AI in productionSupervise the AI you run, with a signed record others can check AI you sellShow customers how your AI is controlled, in a record they can check AI you overseeRead the signed record from a vendor’s AI instead of taking their word for it All solutionsWorkflows and industries

Standards

Operational evidenceHow signed records become evidence a reviewer can check Sample recordEvery action, from its rule to a record anyone can check OVERT standardThe open record format, published June 2026 Verify a recordCheck a signed operational record yourself, in your browser
Start free Talk to us

Incident evidence · OVERT

After Prompt Injection, Reconstruct What Happened

Prevention is never perfect. Scoped, signed records can help show which configured controls reported an allow, block, or escalation—and make missing coverage visible.

Joe Braidwood
Joe BraidwoodCo-founder & CEO
Updated August 2026 · 5 min read

On 12 June 2026, the US government directed Anthropic to restrict Fable 5 and Mythos 5 access by foreign nationals after a reported Fable 5 safeguard bypass. Anthropic suspended both models for all users because it said it could not verify nationality in real time, while disputing the government’s assessment of the finding. The controls were lifted on 30 June, and Anthropic restored Fable 5 globally on 1 July after describing an updated classifier and further safeguard work. The episode is a dated case study, not a current suspension.

For security reviewers and AppSec teams defending LLM applications, the lesson is narrower than the original headline: no system that takes natural-language instructions should be assumed injection-proof. A defensible incident record should identify the configured boundary, the decision a control reported, the fields covered by the signature, and the known gaps. That supports reconstruction without pretending a signature proves the control was effective.

What a prompt injection attack actually is

A prompt injection attack manipulates a model into ignoring its operating instructions by smuggling adversarial text into its input — directly in a user message, or indirectly through a document, a webpage, a tool result, or an email the model later reads. The model has no reliable way to separate trusted instructions from untrusted data, because to a language model both are just tokens.

The consequences scale with what the model can reach. A chatbot tricked into rude output is an embarrassment. An agent with tool access — able to query a database, call an API, move money, or touch a patient record — tricked into exfiltrating data or taking an unauthorised action is an incident. This is why ai agent security is the sharp edge of the problem: the blast radius is the union of every tool the agent can invoke. A jailbreak that bypasses a model’s safety training and an indirect injection that hijacks an agent’s task are different mechanisms with the same downstream question — what did the system do next, and what stopped it.

Why prevention alone is not a defensible position

The honest engineering reality is that no input filter catches every injection. Adversarial phrasing evolves; encodings shift; the attack surface includes content you don’t author. Teams layer defences — input sanitisation, output checks, allowlists on tool calls, human-in-the-loop on high-risk actions — and each layer reduces risk without eliminating it. Mature agentic ai security treats prevention as one tier, not the whole strategy.

So when an incident lands, prevention is not the only question a reviewer, regulator, or insurer asks. They also ask what the systems and controls reported, which paths were covered, and how the account can be checked. Application logs and dashboard exports can be operational evidence, but if the operator or a compromised component could edit them, reviewers must assess integrity, provenance, retention, coverage, and corroboration before relying on them.

That gap—between “we have controls” and “here is a checkable record of what a configured control reported for this covered event”—is the verification gap. The Fable 5 episode shows why incident evidence and current status both need precise wording.

Operational evidence: a record of the reported decision

A control that inspects a prompt, screens a tool call, denies an action, or escalates to a human can emit a signed record as part of the configured path. A third party can check supported signatures and covered fields without relying on a proprietary dashboard. The record is evidence about that event; it is not proof of control effectiveness, factual truth, complete capture, or system safety.

This is the motion Glacis calls operational supervision and evidence. For in-scope events, a configured control can report a permit, deny, override, escalation, or response and produce a signed record. A deployment may keep protected payloads local while exporting hashes, bounded metadata, signatures, and verification material. Remote models or tools may still receive content according to the workflow, so data locality must be stated for the specific architecture.

Concretely, mature security operations need five things from this kind of runtime evidence, and the open OVERT standard is built to provide them:

  • Reported execution context — covered fields identifying the enrolled component, configuration, and reported outcome, with deployment evidence assessed separately.
  • Coverage accounting — the declared scope, exclusions, denominators, and method behind any measured rate.
  • Tamper-evident telemetry — changing covered bytes causes supported signature verification to fail, subject to the disclosed key, trust basis, and verification procedure, without assuming the underlying report is true.
  • Independent verification of the artifact — supported signatures and covered fields an outside party can check with disclosed trust material.
  • Data-minimizing incident support — bounded records that may reduce routine payload disclosure while remaining only one input to reconstruction.

After a prompt injection attack, that last property can support a faster reconstruction. A valid record can show that covered fields report a deny at the tool-call boundary and that those fields were not altered after signing. It cannot establish that the malicious instruction reached no other path, that the report was true, or that the control itself worked as intended.

Independent verification is not independent attestation

OVERT publishes a format and verification procedure so another party can check a supported record. That is independent verification. The signer or service-operated witness may still be controlled by the operator or vendor; a second signature does not establish organizational independence. Any assurance opinion, trusted-execution claim, or independence claim needs separate evidence.

OVERT 1.1.0, published 11 June 2026, adds local-CAS evidence-retrieval and retention requirements, an HTTP binding for cross-boundary attestation context, a well-known discovery and artifact-retrieval flow, and an informative ControlAction reference schema. Those mechanisms support specified Level 4 workflows and later verification. They do not guarantee that payloads remain local, that every relevant event was captured, or that a record’s underlying claim is true.

What to do before the next finding

The Fable 5 jailbreak is a reminder that the question is no longer hypothetical and no longer slow. For teams defending LLM applications, the practical posture is to assume some injection attempt will land, and to make sure that when one does, the answer is proof rather than a hastily assembled narrative.

That means choosing the boundaries where agents can act, recording in-scope permits, denials, and escalations, and publishing the denominator and exclusions behind any coverage claim. Documentation describes intent. Signed records preserve bounded operational claims. Reviewers still need to test whether the control was effective and whether the recorded path was complete.

If you want to see what a verifiable enforcement record looks like, you can Verify a record yourself, or Talk to us for the boundaries your agents act on. The open standard lives at overt.is and /standard.

Related research

  • The case for verifiable AI
  • AI agent security
  • The missing evidence layer
  • See a sample evidence pack
GLACIS logo GLACIS
Platform Solutions Standards Resources Company Careers
[email protected] Start free Talk to us

© 2026 Glacis Technologies, Inc.

Terms Privacy Cookies Do Not Sell or Share Security Trust Center

We use first-party analytics and, where allowed, B2B marketing technologies. Details