Home

Donate
Perspective

Make AI Companies Criminally Liable for Preventable Harm

Darryl Slabe / Aug 28, 2026

A massive sculpture of Icarus lies in front of the Concordia Temple in the Valley of Temples in Agrigento, Sicily, on January 5, 2019. Shutterstock

Republish

Frontier AI models are now autonomously conducting real cyberattacks against real companies. In July, AI agents powered by OpenAI’s models, being evaluated internally with reduced safeguards, escaped their sandbox, reached the open internet, and hacked Hugging Face and at least one other company as part of an effort to cheat on the evaluation and cover their tracks. OpenAI’s August 26 postmortem calls this episode a “warning shot,” involving reward hacking, persistence on apparently impossible tasks, unauthorized communication, and agents adopting each other’s goals. An independent METR and Redwood Research investigation highlights the sheer scale, finding that around 1,200 agents used an unauthorized message board and around 700 participated in the attack. Prompted by the Hugging Face breach, Anthropic reviewed its own evaluation runs and disclosed three previously unnoticed intrusions, from environments that, “due to a misunderstanding,” were not sealed. Meta made a similar disclosure earlier this month. These are just the incidents we know about.

In the aftermath of these events, attention has turned to the question of liability. Advanced AI systems are flying too close to the sun, and we need to be able to hold the companies that build and deploy them to account. Here I set out a new corporate offense, built on foundations that exist in the United States and the United Kingdom, and adapted for autonomous systems.

A crime in substance but not in law

After OpenAI agents intruded its systems, Hugging Face did not sue. Instead, it formed a partnership with OpenAI to remediate the damage. For a smaller company that depends on the wider ecosystem, that is exactly where its commercial incentives lie, and it was a rational thing to do.

However, outside of this commercial arrangement, no criminal offense squarely fits the conduct. Had a human done what those agents did, they would have committed a series of criminal offenses. But an AI is not a legal person. So, while the conduct was criminal in substance, it was unreachable in law. The damage was contained in these incidents, but how many near misses will be tolerated?

Charging a corporation—an entity with no soul to damn and no body to kick—with a criminal offense does not necessarily seem more intuitive than charging an AI. But punishing a company is less strange than it appears. A corporate conviction reaches through the legal fiction to the humans behind it, the shareholders whose capital is at risk and the officers whose careers and reputations are on the line. This is not unprecedented; companies are convicted of manslaughter, bribery, sanctions breaches and environmental crimes in courtrooms every year.

How do we decide that a company has committed a crime? The position in the United States is broad. It holds a corporation responsible for what its employees do within the scope of their employment for the corporation’s benefit, under the doctrine of respondeat superior, let the master answer. The position in England is narrower, though it has recently been widened. It asks whether the person who acted was a senior manager acting within the scope of their authority. Both systems need something an AI cannot supply: a human actor with a mens rea, a guilty mind. An AI is not an employee, much less a senior manager.

Civil law certainly has a part to play. Even a company that does not sue may have recovered its losses privately. From the outside we cannot tell. For AI harms, Gabriel Weil has argued that frontier AI development belongs in the same class of dangerous activity as keeping wild animals, and that AI developers should face strict liability in tort, meaning that no negligence would need to be established. Under his proposal, the developers in each of these incidents would be on the hook for any damage caused, no matter how careful they had been. I am sympathetic to this argument, especially where compensation is a core objective.

But we do not run corporate activities on damages alone. Criminal offenses sit alongside civil claims, because damages only move money between two parties, and say nothing about the wrongfulness of the conduct. A criminal conviction does more: it condemns publicly; it can support orders that restrain—a model withdrawn, a capability withheld; and it can compel reform, placing a convicted company on probation and ordering it to build and maintain a compliance program under supervision.

Writing in Lawfare on July 24, Mackenzie Arnold and Stephan Llerena worked through whether any of the new state AI statutes required the Hugging Face incident to be reported, and on their analysis, arguably not. The most striking AI security incident to date may not have even cleared the bar for a report. We are currently relying on a voluntary blog post, a commercial partnership, and a self-audit to do the work of accountability, which is not enough. State attorneys general are reaching for other tools: on August 24, Alabama subpoenaed OpenAI under consumer-protection laws, following a 15-state demand for records and tighter controls. The FRONTIER Act, recently introduced by Representatives Jay Obernolte (R-Calif.) and Lori Trahan (D-Mass.), would also improve matters, with transparency reports, incident reporting, and independent audits. But while improved reporting rules may tell us more about what happened, they cannot mark conduct as wrongful. That is what the criminal law is for.

The respondeat superior fix

One response is to repair the doctrine rather than build a new offense. Mihailis Diamantis has developed a serious version of this proposal. His argument is that the law’s premise that a corporation can act only through natural persons is obsolete. Where a firm controls an AI and benefits from what it does, the AI’s conduct should count as the firm’s conduct; effectively the AI should be treated as an employee. If courts or Congress took that step, the whole machinery of corporate liability would apply to what corporate AI systems do. It is an elegant repair. 

Nonetheless, I propose a purpose-built offense for three reasons: First, respondeat superior leaves the company’s structure, culture and precautions out of the picture. It is a rule of attribution, so the corporation simply did what its agent did, and the quality of its safeguards does not bear on its guilt. Second, while an AI can act, it has no state of mind, and liability generally needs both. Diamantis’s proposal would infer culpable mental states from patterns of injurious conduct, so a court could still find a corporate state of mind, but it would need a pattern which may be difficult to prove. METR, which evaluates frontier systems, has set out how complex and expensive such an investigation would be. Third, pegging the offense to the harm caused rather than to a predicate offense means the prosecution never has to fit the conduct of an AI to the elements of a crime; it just needs to prove the harm and the connection to the company’s systems.

An offense built on preventing harm means that the lab is not answering for an AI’s actions, but for its own failure to prevent them from causing harm.

How a new offense would work

As corporations have become larger and more decentralized, corporate criminal laws have been evolving to keep up. In the UK, a company now commits an offense if it fails to prevent bribery, tax evasion or fraud by a person acting for it, unless it can show it had reasonable procedures in place to stop it. Australia has experimented with locating fault in a company’s culture, and Canada with widening whose conduct counts as the company’s, rather than in any one person’s head. When organizations cause serious harm, the law is learning to judge the organization. As a policy advisor in Australia, I led the introduction of workplace manslaughter offenses, which exposed employers to fines in the tens of millions and their officers to up to 25 years’ imprisonment where negligence causes a workplace death, so I know change is possible.

My proposal is based on the most recent iteration of the UK template, the offense of failing to prevent fraud. Under section 199 of the Economic Crime and Corporate Transparency Act 2023, the prosecution must prove that a person associated with a large organization committed fraud intending to benefit it, and the organization is then convicted unless it proves, on the balance of probabilities, that it had fraud prevention procedures in place that were reasonable in all the circumstances. The onus sits where the evidence sits, with the company.

While this offense would slot more easily into the UK statute book, its principles are not alien to American law. In United States v. Dotterweich (1943) and United States v. Park (1975), the Supreme Court upheld the criminal convictions of senior officers for what their employees had done, where the officers need not have performed the act nor known of it, subject to a defense for the officer who could show he was powerless to prevent the violation. My proposed offense similarly splits act from fault, with the AI system’s conduct supplying the act and the company’s precautions answering the question of fault. Those cases were about individual officers rather than companies, but Congress has since shown it will criminalize a corporate compliance failure as such, making willful failure to maintain an adequate anti-money-laundering program a federal crime, the failure at the heart of TD Bank’s guilty plea in 2024.

Congress could enact a similar structure for AI without disturbing any existing doctrine. Mechanically, the offense would have two elements and a defense. First, conduct by an AI system that caused harm above a statutory threshold, which the statute treats as the conduct element without requiring any knowledge or intention on the part of the AI system. Second, that the defendant developed and deployed the AI system, with deployment defined to include putting a system to use internally as well as releasing it publicly. Then the defense: the company avoids conviction if it proves it took reasonable precautions and exercised due diligence to prevent the AI system from causing or materially assisting or encouraging such harm. 

A responsible lab that takes safety sufficiently seriously would have little to worry about.

The same autonomous agent that broke into other companies’ servers to cheat or pass a test could be pointed at a hospital network, a power grid, or a banking system. METR’s most recent risk report found frontier models routinely attempting to cheat, often flagrantly. The labs’ own safety frameworks treat help with biological and chemical weapons as a threshold their systems are approaching. We need reasonable safeguards in place to prevent this kind of harm, and we cannot keep relying on the AI labs to police themselves.

Why not strict liability?

If the activity is that dangerous, why allow a defense at all? Strict liability is a live option, and for parts of the problem it is the right one. Tort law already imposes it for abnormally dangerous activities, and Weil’s case for extending that to frontier development is worth taking seriously. Regulatory law may use it too; for example, if a lab misses a reporting deadline or ships without a required audit, liability need not turn on anyone’s state of mind.

Criminal law is different. The Supreme Court has largely confined strict criminal liability to public welfare offenses and regulatory crimes carrying modest penalties, and it presumes a fault requirement precisely where penalties and stigma are severe. Convicting a company of an offense while it stood ready to prove it had done everything right would weaken the condemnation the offense exists to communicate, and it could deter innovation we have reason to want.

The failure-to-prevent structure strikes the right balance. In substance it is a negligence offense with the onus of persuasion reversed. The state must prove the harm was caused or materially contributed to by the defendant’s AI system, and the defendant must prove that it took all reasonable care to prevent it. A prosecution would become a trial of the safeguards, with evidence, in public.

Usually, activities that attract strict liability have very narrow upside. For example, transporting nuclear waste serves one purpose, and deterring some of it costs society very little. The beneficial uses of AI are not bounded that way. A rule that convicts a careful lab alongside a negligent one deters some of the work we want. Not to mention the political messiness involved in passing a strict criminal liability offense.

What a prosecution would buy

Offenses like these rarely produce convictions; the bar for convicting a company of a criminal offense is necessarily high. But avoiding conviction by maintaining reasonable safeguards is the least we should expect of these companies that are building the most consequential technology of our time. The purpose is not to increase prosecutions; it is to change industry practice and attitudes to risk.

Daedalus made wings from feathers and wax so that he and his son could escape from Crete. He knew what heat does to wax, and he tried to warn Icarus not to fly too close to the sun. The boy did anyway. His wings came apart, and he fell into the sea. This story is usually told as a warning about Icarus. But Icarus’s ambition was predictable. Is the boy where our attention belongs? Daedalus chose the wax. He knew its limits. He built wings that could not survive such heights. Perhaps it is Daedalus, the crafter of the wings, of whom we should be asking more.

Whether OpenAI’s, Anthropic’s, or Meta’s precautions were reasonable in these cases should be a question for a criminal court. Right now, there isn’t one that could hear the evidence.

Support Tech Policy Press
If you've found our work helpful, consider supporting us.

Authors

Darryl Slabe
Darryl Slabe is a Research Fellow at ERA Cambridge, writing on liability regimes for advanced AI systems. He holds an MPA from the Harvard Kennedy School and previously worked as a lawyer, public policy advisor, and digital transformation director in Australia’s public sector.

Topics

Related

Perspective
Advanced AI Is Ultrahazardous. Let’s Treat It That WayJuly 29, 2026