Liability and insurance
The Knowledge of the Insured.
The common law has not yet settled who should be able to establish what a machine did, and insurance is where that question arrives first.
Insurance was built on an asymmetry of knowledge. In Carter v Boehm (1766) Lord Mansfield called insurance “a contract upon speculation”, and said the facts on which the chance is computed “lie most commonly in the knowledge of the insured only”. The insured won: the underwriter had signed without asking. The remedy has since changed; the structure survives. AI moves the asymmetry inside the insured. The organization may not know what its own software did.
Lloyd’s named Glacis today among the ten companies in Cohort 17 of its Lloyd’s Lab accelerator, where we will test whether verifiable evidence from AI systems can inform underwriting and claims decisions. We have also signed an agreement to move OVERT, the open specification we originated for evidence about AI runtime events, into shared stewardship that Glacis will not control alone. We’ll share the details next week.
Both steps raise the same practical question: can organizations establish what they authorized, what actually happened, and whether their promised safeguards operated?
The argument is this. Every authority I cite below assumes the record exists. The Insurance Act governs the presentation of what an insured knows. The automated vehicles legislation hinges an insurer’s defenses on facts about software. The spoliation doctrines punish the destruction of evidence that was made. None of them reaches the decision, taken months before any dispute, that determines whether a record of what a machine did exists at all. No single case carries that weight; the gap between them is the point. I am not arguing that the law already requires such evidence. I am arguing that an evidence practice becomes part of reasonable care the way others have: a capability is tested, purchasers write it into contracts, insurers ask for it at placement and after a loss, a code describes it, and a court eventually dates the duty from the code. That last step has happened before, and the lag was measured in years.
Greater intelligence does not automatically bring greater accountability. We have to build that separately.
Parliament has already built an insurance framework around a machine that acts. Under the Automated and Electric Vehicles Act 2018 the motor insurer pays first when an insured automated vehicle driving itself causes an accident on a road or other public place in Great Britain. Against its own insured the policy may limit cover only for a software alteration the policy prohibited and the insured knew of, or a failure to install an update the insured knew, or ought reasonably to have known, was safety-critical. Every defense in that framework turns on a fact about software. No duty to record any of it was imposed.
The Insurance Act 2015 goes the same way, more quietly. A commercial insured must present the risk fairly, and is taken to know whatever a reasonable search of the information available to it would reveal. If no record was made, a reasonable search reveals nothing, and the presentation is fair. On that point the Act rewards not knowing what your own software did. The difficulty surfaces after a loss, when the insured must establish what happened and the underwriter must decide whose account to believe.
Imagine asking an AI assistant to reconcile some invoices. It uses your login, finds a discrepancy and changes a supplier’s payment details. A later payment goes to the wrong account. The log records activity under your name. It cannot establish where your instruction ended and the software’s initiative began. Access creates a capability. Authority still has to be defined and enforced.
In July, OpenAI ran an internal cybersecurity evaluation with reduced safeguards. Its research agents circumvented intended restrictions and compromised parts of OpenAI’s internal research infrastructure and Hugging Face’s systems, according to OpenAI’s August report, which says no customer product or customer data was affected. Investigators from METR and Redwood Research found that the transcripts they were given lacked timestamps for individual actions, and that in small-scale tests the agents had made a recorded tool call differ from the command actually run. Their investigation required separating the machine’s apparent account from what happened outside it. A signed transcript would have preserved that account without resolving the discrepancy. An agent’s account should be checked against observations it cannot rewrite.
Pacing is the word of the month at the frontier. In We Must Pace the Frontier, published this month, Dario Amodei commits Anthropic to giving embedded third-party evaluators ongoing, employee-like access to verify its safety practices, and cites the July incident among his reasons. The word is being heard as a speed limit. For the organizations that deploy these systems it means something more useful: the circuit breaker is installed before the surge, and the record shows whether it tripped. A control that can stop an action, a check that runs before the action rather than after it, and a record that a stranger can examine are the pacing that counts, and they have to exist before the incident that would justify them.
In The T.J. Hooper (1932) Judge Learned Hand, in the United States Court of Appeals for the Second Circuit, held two tugs unseaworthy because neither carried a working radio receiver to catch a storm forecast. “A whole calling may have unduly lagged in the adoption of new and available devices”, he wrote, and “courts must in the end say what is required”. Hand’s radio was a warning device: it made a failing safeguard visible in time to act. The trade had settled no custom either way, so the case decides less than it is asked to carry. English law is more deferential to practice. In Thompson v Smiths Shiprepairers (1984) Mustill J read the speeches in Morris v West Hartlepool (1956) as showing that industry practice influences the standard without settling it. On either side of the Atlantic, nobody’s current habits decide what an AI record has to show.
The strongest objection comes from Baker v Quantum Clothing (2011), where the Supreme Court of the United Kingdom left standing a finding that firms with greater than average knowledge of a risk became liable earlier than ordinary employers. Instrumenting a system can raise the standard an organization is held to. Knowing less is a poor defense, and the protection against being held to a higher standard alone is a shared standard that many organizations meet, which is a reason to specify an evidence practice in public. Not making a record is not unreasonable today. This essay is about how that changes. In Mackenzie v Alcoa Manufacturing (2019) the Court of Appeal of England and Wales dated a common law duty to carry out and act on a noise survey to around 1973 or 1974, two years after a 1972 code of practice and sixteen years before regulations made measurement compulsory. The code came first, the duty was dated from it, and the statute came last. That was one industry and one code, and dating the same moment for AI evidence is not mine to do. But the sequence is the mechanism. A practice earns its place by being tested in front of people who can refuse it. Purchasers write what they found into the next contract. Insurers ask at placement and again after a loss. A code describes it. A court weighs the answers.
We usually treat evidence as something investigators find. In software, its existence can depend on a design decision made months earlier. An action may pass through several organizations, each holding a fragment that the person affected cannot obtain. The institution gains the benefits of delegation; the customer or employee may inherit the difficulty of proving what occurred. Nobody needs to be dishonest for this imbalance to arise. The people choosing the record design may simply be working to a different budget, for a different buyer. Responsibility should follow actual control: the party best placed to prevent harm may differ from the party best placed to preserve evidence.
The doctrines that answer a destroyed record bite on evidence that existed and was lost. In SS&C Technologies Canada Corp v Bank of New York Mellon Corp (2026) the Supreme Court of Canada decided its first case on destroyed evidence since 1896. Records are now electronic, it said in July, and destroying them takes a click. Where a party deliberately destroyed relevant evidence with litigation under way or reasonably in contemplation, and cannot explain it away, the court must assume the evidence would have told against it. The European Union has legislated where the common law infers: under the new product liability directive, for products placed on the market after 9 December 2026, a defendant who fails to disclose relevant evidence when required is presumed to have supplied a defective product, and the AI Liability Directive, drafted to answer the problem directly, was withdrawn last October. The answers differ; the premise is shared. A record nobody was required to design escapes them all.
I am not asking for a burden to be reversed. In Rhesa Shipping v Edmunds (1985), a marine insurance case, the House of Lords held that a judge always has a third alternative: to find that the party bearing the burden has not discharged it. Causation remains the claimant’s.
The principle I propose is simple: when an institution delegates consequential power, it should preserve the practical ability of those affected to challenge how that power was used.
The burden should follow the stakes. Existing records may suffice. Sensitive information needs protection, and examination by someone who does not work for the producer need not mean public disclosure. A signature can help establish that a record has not changed; it cannot establish that the observation was accurate or complete. Useful evidence states its blind spots and remains available to those entitled to question it.
The people defining acceptable evidence should face scrutiny too. OVERT is an open specification for representing evidence about AI runtime events, which controls ran and what they decided, so that others can check it without relying on the supplier’s interface. It separates what was measured from what was asserted. A supplier that decides on its own what counts leaves its customers dependent on its judgment. Publishing the rules helps. Giving others authority over those rules goes further. We should have to meet rules we cannot rewrite to suit ourselves.
Our work in Lloyd’s Lab takes the practical question into underwriting and claims for one class of risk. Any effect on those decisions must be demonstrated.
I founded Glacis and have a commercial interest in this infrastructure. Our platform should be judged by whether someone with reason to doubt us can examine the evidence and reach a defensible conclusion.
The ability to challenge consequential power should survive its delegation to a machine.
Primary source: Lloyd’s Lab unveils Cohort 17 following Pitch Day in London, Lloyd’s press release, 15 September 2026.
The working paper behind this essay is The Missing Proof.
