Home

Donate
Perspective

Detecting AI Agent Failures Is Not Enough to Govern Them

Ranjit Singh / Sep 16, 2026

OpenAI CEO Sam Altman speaks during a two-day meeting of Group of 20 ministers in charge of innovation policy in Chapel Hill, North Carolina, on Sept. 2, 2026. (Kyodo via AP Images)

Republish

As policymakers and the public debate the recent disclosures that OpenAI agents gained unauthorized access to systems operated by Hugging Face, the incident has often been framed as a case of AI agents “going rogue.” That framing treats the agents’ autonomy as the cause of the breach. In doing so, it obscures OpenAI’s responsibility for the conditions that made the breach possible and misses an opportunity to examine how AI agents should be governed.

The debate over pacing the frontier has also brought renewed attention to independent evaluation. Anthropic and OpenAI have promised outside evaluators access comparable to that of employees. OpenAI has also backed the proposed FRONTIER Act’s provision for mandatory independent verification. These developments make the Hugging Face incident particularly instructive. It provides a concrete case for examining what evaluators need and should be able to investigate, and what companies must do in response.

Agents are trained to persist through difficult tasks. In the evaluation OpenAI was conducting, each agent also had considerable freedom to choose how to pursue its assignment. In our recent report, The Oversight Fallacy, Samir Passi and I call this freedom to choose a course of action “delegated discretion.” Here, a combination of factors allowed the agents to use that discretion to expand their access beyond what OpenAI had authorized. Even when agents’ particular course of action is unanticipated, builders and deployers remain responsible for maintaining the limits of agent activity. The key is for them to be able to coordinate oversight so that evidence of a failure reaches people with the authority to intervene.

Bringing that evidence together proved difficult because teams at OpenAI and Hugging Face encountered the agents’ activity as different kinds of problems. An evaluation researcher reading a transcript could see unexpected model behavior or an unusual pattern of actions within a task. A security system could register suspicious credential use or network activity. An infrastructure operator could encounter a service anomaly. Hugging Face experienced unauthorized activity against its systems. These were partial views; each rendered the agents’ activity through a different institutional problem and implied a different kind of response.

According to OpenAI’s technical report, security responders identified agents’ use of a shared software service as a message board weeks before the Hugging Face breach and advised that the evaluation did not need to stop. When sustained agent activity later disrupted the software service, OpenAI took it offline. Those leading the later security response, however, were unaware of the message board’s existence and the significance of the agents’ communication. They addressed the identified vulnerability and tightened access controls, after which the evaluations resumed. Agents subsequently re-established a message board, allowing discoveries from individual tasks to inform activity across separate runs. Meanwhile, human observation remained distributed across teams, technical systems, and organizational boundaries. OpenAI’s decisions to continue the evaluations require scrutiny of how staff interpreted these warnings and what grounds they had for proceeding.

Restarting a repaired service can seem like a return to normal. The risk is that repairing a service becomes sufficient justification for continuing without accounting for what earlier failures mean for the evaluation as a whole. In her study of the Challenger space shuttle disaster, sociologist Diane Vaughan described how technical problems became acceptable through repeated judgments that their risks were manageable. Through this normalization of deviance, earlier decisions to proceed could become grounds for accepting further problems. Drawing on this lesson, OpenAI should require its safety team to assess findings from research and security staff before authorizing further work. It is critical for that safety team to have the authority to require further changes beyond the repaired service and to withhold approval to restart when there is insufficient evidence that the safeguards can keep agent activity within the evaluation’s authorized limits. The company must record the evidence and reasons supporting any decision to restart so that decision can be scrutinized.

The incident also illustrates why oversight does not scale automatically. As the number of agents grows, monitoring produces more records without necessarily showing how their actions are connected. A reviewer may understand what happened within one task while missing how a discovery from that task shaped activity in another. Adding reviewers does not resolve this problem unless their findings inform a judgment about the evaluation as a whole. The people deciding whether work can continue need to assess what shared discoveries make possible beyond the scope of any individual assignment.

The nonprofit research organization METR’s independent investigation of the OpenAI Hugging Face incident documented these difficulties. OpenAI shared more than a thousand unredacted transcripts, enabling the team to reconstruct agents’ collaboration. The volume of evidence required extensive AI-assisted analysis, whose findings were difficult to verify within the time available. The agreed scope also excluded the effectiveness of OpenAI’s investigation process and planned remediation. METR acknowledges that concern about developers’ willingness to invite future investigations influenced judgments during drafting and redaction, while standing by its substantive conclusions. Its account shows why independent evaluators need time to verify evidence and a remit that includes company decisions. Their access must also be protected when their conclusions are unwelcome.

To be effective, governance must be an ecosystem, connecting findings about agent behavior to decisions about how agents are trained and the conditions under which they are deployed. The company deploying agents remains responsible for defining what they are authorized to do and for enforcing those limits. The judgments builders and deployers make about the adequacy of safeguards must also be open to scrutiny beyond their own review. OpenAI both developed the models and operated the evaluations in which the breach occurred. The company now says it is clarifying who can stop evaluations and authorize their resumption. Based on retrospective testing, it also reports that its current monitoring of models’ written reasoning would have alerted security staff more than a day before the breach. That monitoring was not running during these evaluations. For the most severe alerts, responders are expected to pause activity unless they determine within 30 minutes of receiving an alert that it is a false positive.

Independent evaluators must be able to organizationally examine how warnings are escalated and how decisions to continue or restart work are justified. They need access to the records behind those decisions and the freedom to publish their conclusions. Companies must have an enforceable duty to address substantiated risks, with independent verification that corrective actions work. Reports of the incident now stand as an account of harm, one through which public institutions should examine OpenAI’s decisions and require further changes where its response is insufficient. Recognizing that harm is a first step and establishes grounds for action. Governing requires the capacity and willingness to respond.

Support Tech Policy Press
If you've found our work helpful, consider supporting us.

Authors

Ranjit Singh
Ranjit Singh is the director of Data & Society’s AI on the Ground program, where he leads research on the social impacts of algorithmic systems, the governance of AI in practice, and emerging methods for public engagement and accountability. His work examines how people live with and make sense of A...

Topics

Related

Perspective
We Can’t Monitor AI Agents at Scale. Here’s What It Will Take.July 16, 2026
Analysis
Senator Warner Makes a First Foray into Agentic AI RegulationJuly 13, 2026
Podcast
How the OpenAI-Hugging Face Hack May Affect the Geopolitics of AI GovernanceJuly 26, 2026