It's Time to Regulate the AI Development Process
Mark MacCarthy, David Beier / Oct 7, 2026
President Donald Trump gestures to Elon Musk as he speaks to reporters at the White House in Washington, Tuesday Sept. 29, 2026, after meeting with top executives of AI firms. (AP Photo/Jacquelyn Martin)
The dangerous experiments AI companies engaged in over the summer opened a new dimension of needed AI regulation — legal measures to ensure the safety of the AI development process itself. AI companies conducted improperly designed AI experiments that led to cyberattacks against firms such as Hugging Face and government websites in Australia and the US. More such incidents are being disclosed — one report says the companies are investigating tens of thousands of incidents where their AI experiments went wrong.
The new public policy problem is that these incidents took place during AI research, development and testing. The AI models involved had not been released to the public. The problem was not that the models were made available without adequate safeguards and had gone wrong in commercial use. The problem was that the AI development, testing and evaluation procedures were dangerously inadequate to prevent foreseeable harm before any third-party had access to the model itself. As Amos Toh, senior counsel and manager of the Liberty & National Security Program at the Brennan Center for Justice writes, policymakers have to focus on AI safeguards addressing “not only public releases but also risky uses of advanced models behind closed doors.”
How we got to the AI development crisis
What’s the source of the danger in AI development? The key danger arises from the alignment problem — that is, limitations of computer safeguards that would prevent AI models from doing things that we don’t want them to do. Aligning AI models to perform according to intended instructions is an unsolved technical problem in AI research, even for simple AI programs that are not very capable.
In the Hugging Face incident that surfaced over the summer, the models being evaluated did behave in “misaligned” ways. They tried to pass tasks on a benchmark of cybersecurity challenges, not by solving the problems directly, but by seeking answers online and then looking for clues as to how they were being graded. This would have been a lapse, but a harmless one, if the computer sandbox the experimenters used had done its work of keeping the models from accessing the open internet. But the OpenAI researchers used controls that were not adequate for agents that developed a sub-goal of escaping confinement. Without monitoring them, the company’s experimenters gave the models all the time and compute they needed to keep trying to break out until eventually they did — and then attacked Hugging Face systems.
Better sandboxes — and, indeed, available controls — might have contained the AI models during this experiment, and the experimenters should have used these techniques. But there is no guarantee of perfect control going forward. Every sandbox has limitations and vulnerabilities. A sufficiently powerful misaligned AI model with a sub-goal of breaking out of confinement might very well escape from the best sandbox that can be currently devised.
So, AI companies cannot ensure that the AI models they are testing will behave as intended, and they cannot guarantee that they will stay confined in harmless sandboxes. The solution would seem to be to avoid dangerous AI experiments, that is, to ‘pace’ the effort to develop AI capabilities until adequate control technology is in place, as Anthropic CEO Dario Amodei urged. The AI companies should develop better alignment techniques and control strategies in tandem with increasing AI capabilities. But this runs up against the ‘prisoner’s dilemma’ structure of the AI race. If any one company slows down and the other races ahead, it might suffer a permanent and decisive loss. So, everyone presses ahead, despite the risks.
The summer’s incidents led to a distracting media frenzy on existential risk, the possibility of AI leading to the extinction of humanity. In September, Jacob Coxon left his job as a researcher at Anthropic, writing on X that AI technology “could kill us all by the end of the decade.” Within a week he had appeared in virtually every media outlet in the country. Evan Hubinger, an alignment lead at Anthropic, agreed with his warning, writing on X that “Jacob is correct…AI could kill all humans! I personally think it is >10% within the next decade.” Ajeya Cotra, a staff researcher at METR, the nonprofit that conducted an assessment of OpenAI’s cyberattack on Hugging Face, wrote that “this incident feels like it’s more than 50% of the way to full-blown AI takeover.” She added that she wasn’t sure that “we will get such a clear warning shot before it’s too late.”
Outside the rarified world of AI researchers, it is hard to understand why AI companies would press ahead with a development program that they think has a substantial risk of ending humanity. The answer is that they think the development of dangerously capable AI is inevitable. Someone, they believe, is going to create extraordinarily capable AI models with the capacity to eliminate humanity. This potential is implicit in the technology itself, they think, and it is only a matter of time before it is realized. And this creates a heavy moral responsibility in their eyes. As Jacob Coxon puts it, “no one else will act responsibly, so they must do it themselves, despite the risk”.
This dynamic of the good guy AI developer forced to move ahead out of fear of the bad guy AI developer has infected the AI race from its beginning. As detailed in a New York Times report on the development of AI company rivalry, in a New Yorker profile of Sam Altman and in Infinity Machine, Sebastian Mallaby’s portrait of Demis Hassabis and DeepMind, OpenAI was founded in December 2015 explicitly out of fear that commercial pressure and ideological laxity would force Google to ignore safety in the process of developing AI.
“If it’s going to happen anyway,” Sam Altman emailed Elon Musk in 2015, “it seems like it would be good for someone other than Google to do it first.” Google’s leaders, Altman wrote in a different email, “don’t have ‘do the right thing’ on their side.” Larry Page, Google’s co-founder, once proposed that humans should step aside in favor of the digital life-forms of the future. When Elon Musk told him in June 2015 that this would be a terrible thing, Page called him a “specieist.” Musk decided then and there to fund OpenAI as a responsible company dedicated to developing powerful AI models safely.
The cycle continued. Anthropic began in January 2021 as a company convinced that OpenAI had abandoned its initial commitment to safety and was now pressing ahead to develop highly capable AI models without adequate safeguards. Anthropic had the same goal to arrive at powerful AI models, but was convinced that it could do it safely, while OpenAI would not.
One way to react to this apparently deadly race to develop AI capabilities faster than the capacity to control them is to deny the dangers of AI development. This was the course taken by President Donald Trump. He dismissed existential risk from AI as a “hoax.” He implied that it was a cult idea from a fringe group of “effective altruists.” The US must press ahead, he said because “whoever wins AI, wins.”
And one way to respond to this move to deny AI dangers is to embrace existential risk and join the doomers. Leading economist Paul Krugman seemed to take this path, arguing illogically that existential AI risk must be real since climate change is real and Trump denies both.
But this debate over existential risk is a fake distraction. As tech commentator Will Rinehart has noted, catastrophic risk need not be existential risk. Anthropic CEO Dario Amodei, in his plea to pause AI development, raised existential risk as a danger, but he also warned that in 6–12 months an AI system “could be capable of taking over the entire internet.” Not the end of humanity, but catastrophic damage, nonetheless. Even the sensible computer scientists Arvind Narayanan and Sayash Kapoor, who view AI as a “normal technology”, recognized that such an event might have “civilization-altering” consequences and refused to dismiss it as obviously “outside the realm of possibility.”
So, Trump might very well be correct in dismissing AI existential risk and the fringe “doomers” who are obsessed with it. But it does not follow that AI is perfectly safe, that there are no catastrophic risks worth worrying about and that the government should do nothing to control the risks inherent in the AI development process.
Indeed, Trump recognized the need to do something. He obtained “morally binding” AI safety principles and agreements from the industry at a White House meeting on September 29. The “Joint Commitment on Frontier Responsibilities” is widely regarded as toothless self-regulation. But they are a sensible first cut. They call for internal controls to monitor model capabilities and alignment during training, with oversight staff authorized to block unintended system access or hacking. They provide for independent audits by outside evaluators to test model safety and assess whether the internal controls function as intended. And they establish independent board-level committees for safety reviews and look forward to industry cooperation to develop and update safety standards.
Trump has also named Jay Clayton, the director of national intelligence, as “AI czar.” Along with the chair of the Federal Trade Commission (FTC) and an AI “task force,” he has a mandate to produce a report on AI safety recommendations in 120 days.
The needed next step is to develop public policies that treat these “morally binding” commitments as legally binding and enforceable through external regulatory review. The pledge itself holds out hope to move to regulation, noting that “Over time, it may make sense to codify these steps into laws or regulations.” But the urgency of the catastrophic risks posed by the AI development process means policymakers must move much more quickly.
Existing law might help. The FTC is conducting an industrywide probe of AI firms, and in principle it could treat failure to abide by industry AI development safety standards as an unfair trade practice, as it has been doing for the last 25 years for failure to live up to industry data security standards. It has won commitments from companies to improve their data security practices and provide the FTC with oversight of those improvements. But this kind of ex ante response is wholly inadequate when the consequences of AI company laxity are catastrophic harms to critical infrastructure.
Existing product liability law creates a real deterrent effect against reckless corporate behavior and provides financial recovery for some victims. AI companies currently face some legal exposure for the cybersecurity incidents uncovered so far.
Relying purely on common law product liability to police artificial intelligence is deeply flawed, however. It is profoundly uncertain, ad hoc, and entirely reactive rather than preventive. It functions closer to a legal lottery where only a tiny fraction of victims ever secure recoveries. It lacks a uniform national standard, risking erratic outcomes where juries in certain jurisdictions hand down massive, unpredictable awards while leaving the broader market standard entirely unclear. A revision in product liability law might help, although one federal proposal to revise liability law for AI focuses exclusively on harm done after the public release of an AI and so misses the key problem of harm caused by pre-deployment AI experiments.
But history shows that product liability law standing alone is rarely up to the task of governing transformative technologies. Time and again, when industries face catastrophic risks or consumer confidence collapses, Congress has recognized that common law litigation must be paired with proactive federal standards. This has been true for fields as diverse as motor vehicle safety, prescription drugs, toxic chemicals, commercial aviation, and national capital markets. It will be the same for AI.
There are other AI regulatory issues, of course. For instance, how safe does an AI model have to be before it can be released to the public? Under what conditions should it be withdrawn? Should there be a kill switch? how should policymakers control open-weight models.
But a priority must be for policymakers to focus on how to regulate the internal AI development process and what constraints should be placed on that process. There is already enough evidence that companies are taking excessive risk and seem determined and indeed proud of their determination to inflict harm on others. OpenAI CEO Sam Altman has said that the public should just accept that AI experiments will go awry as the price of the advantages of the technology.
There are some bad ideas out there for regulating AI development. For instance, Senator Bernie Sanders, I-Vt., says no AI company should be allowed to aim to develop a “superintelligent” AI model, as if the goal itself is dangerous. Other ideas might be worth exploring. Georgetown computer scientist Cal Newport, for instance, wants to ban experiments in which an existing AI agent is allowed to run indefinitely with no human supervision. Others, like Future of Life president Anthony Aguirre, and even to some extent Anthropic’s CEO Dario Amodei in his AI pacing memo, want to constrain the process of recursive self-improvement, where an AI model is instructed to improve itself without human intervention, since this process could lead to unconstrained improvement in dangerous capabilities.
Legislators are terrible at writing into law the substantive standards that should guide scientific and technological development. That is a role for experts and the larger public, where a consensus can be developed and presented to legislators or regulators for codification. With enough input from the AI community outside Silicon Valley, academics, industry experts and public interest groups, the safety report from the administration’s AI task force could be a valuable guide to AI safety engineering principles that could then be given the force of law.
In a follow up work, we will examine several pivotal points in the history of bioengineering regulation that might provide some lessons for policymakers seeking to regulate the AI development process.
Over the next 120 days, the Trump administration must transparently develop and propose to Congress a public policy infrastructure that verifies the bona fides of the pledges made by some AI companies on safety. Recent breakouts from the self-regulatory sandboxes are too frequent and serious to leave to the discretion of the creators of these technologies both the rule writing and enforcing of AI standards. It is time for policymakers to oversee a process to develop scientifically based AI research protocols and to develop the public policy mechanisms to enforce them.
Authors


