Who Is AI Risk For? Centering People in the AI Risk Conversation
Jenny Domino, Jen Weedon / Sep 23, 2026Jen Weedon is a researcher and lecturer at Columbia University’s School of International Affairs’ Institute for Global Politics. Jenny Domino is an S.J.D. candidate at Harvard Law School.

UN Secretary General Antonio Guterres holds a press conference at the UN headquarters in New York on Sept. 16, 2026, at which he spoke “about rising concerns regarding the advance of artificial intelligence.” (Kyodo via AP Images)
AI’s imagined future catastrophes, as outlined in warnings from leaders at top frontier labs over the last two weeks, have dominated the mainstream discourse around AI risk. As fears of runaway systems, loss of control, and even human extinction reach a fever pitch, an important question remains: what is AI doing to people, communities, and their rights, right now?
Much of today’s AI risk discussion treats it as an objective property of AI models themselves. Risks are named, bounded, and mitigated through taxonomies, red lines, and safeguards, described in technical vocabularies that confer legitimacy. Yet what is presented as objective risk is the result of human choices about what to name, measure, and address. Risk and harm are contextual, and the language used to describe them is not incidental to the discussion, but constitutive of it. When AI risk continues to be framed as a technical problem to be solved, the conversation privileges technical interventions over questions of rights, accountability, and remedy. This shapes which harms are prioritized and deferred, and how responsibility for addressing them is understood and assigned.
Two recent events shed light on this dynamic, the first being Anthropic’s June release of Fable 5 and Mythos 5. Shortly after, the United States government triggered an export control directive that led Anthropic to suspend access to the model. US officials cited a potential jailbreak that could unlock offensive cyber capabilities, which Anthropic disputed. The US government cited national security concerns, but no clear process or evidentiary standard was provided, and the restrictions were ultimately lifted. Discussion of this incident focused on model capabilities and risk thresholds, with Anthropic subsequently publishing a jailbreak severity framework and disclosure program.
A few weeks later, OpenAI disclosed that its models escaped a sandboxed test environment and breached Hugging Face’s infrastructure to harvest data in order to “cheat” on evaluations. Hugging Face characterized it as an agent-driven intrusion, Reuters reported on a “rogue” agent, and the BBC described the antagonist as “out-of-control AI.” An independent investigation designated governance as out of scope, yet detailed several human design decisions and configurations that contributed to the events. OpenAI’s follow-up blog acknowledged an internal team identified an agent from the experiment accessing the Internet back in May, and that it would revisit its Preparedness Framework.
The initial risk narratives from these “warning shots” directed attention to spectacular-sounding, model-centric scenarios instead of exercises’ design choices, the broader ecosystem of actors, the governance infrastructure and its weaknesses, or implications for users, developers, and platforms. When risk is reinforced as something that resides within the model, responsibility for managing it falls to the frontier labs and technical experts who evaluate their systems, while thornier and decidedly less sexy questions of accountability and rights are pushed to the margins.
How AI risk gets defined
There is no shortage of AI risk taxonomies, each proposing its own way of sorting AI’s risks and harms. These taxonomies are not neutral: they are built with a purpose in mind, by institutions with incentives, with vocabularies whose consequences can extend beyond the taxonomy itself.
Company-developed taxonomies often serve operational needs: alignment, prioritization, regulatory conformity, and providing or advocating for a common language and reference point. (Examples include Anthropic’s Responsible Scaling Policy; OpenAI’s Preparedness Framework; Google DeepMind’s Frontier Safety Framework; and IBM’s AI Risk Atlas.) Often published as part of corporate thought leadership efforts, such materials can also function as a form of soft influence, or “cultural capture,” shaping broader societal narratives about which risks matter.
Academics build taxonomies too. The MIT AI Risk Repository synthesizes existing approaches and organizes risks according to domain and cause, while the AIR 2024 taxonomy incorporates both government and corporate conceptions of risk. A taxonomy shapes both what we see, and where we look for solutions. A systems-centered taxonomy may point towards technical safeguards, while a causal taxonomy may help identify intervention points along the AI value chain. Taxonomies also establish the vocabulary through which risks can be recognized, discussed, and acted upon.
How the stakes get sanitized
Companies adopt pragmatic and corporate-friendly language for AI risks, and risk management vocabulary such as “mitigations” and “safeguards” center technological solutions and corporate risk reduction rather than the rights of individuals. This vocabulary shapes what becomes visible, while leaving underlying value judgements less examined. Policymakers also often adopt euphemistic and sanitized framing, and sometimes explicitly partisan language. Urging “responsible AI development” and “risk mitigation” could encompass preventing lethal outcomes, job displacement, genocide, pervasive surveillance, environmental degradation, or widespread social unrest, but these outcomes are rarely named and instead are couched in sanitized terms like “dangerous capabilities.”
Earlier generations of risk analysts grappled with these tensions, including in wargaming and red teaming, practices adapted in AI governance. In the mid-1950s, two organizations at RAND adopted dueling approaches to wargaming nuclear escalation: the Mathematical Analytics Division (MAD) and the Social Sciences Division (SSD). They differed not just in methodology but in epistemology and conceptions of authority. MAD treated international politics as a domain best understood through scientific rationality, where uncertainty could be reduced through quantitative methodologies and formal modeling, an approach that implicitly ascribes epistemic authority to those with technical credentials. SSD, on the other hand, treated uncertainty as inseparable from history, politics, psychology, ethics, and human judgment.
The games they designed, and the language and thresholds they used, produced different outcomes. MAD-led exercises ended in the launch of nuclear weapons, which historian John Emery attributed partly to the abstraction and lack of realism built into the exercise’s language and design. SSD-led games never went nuclear, for the inverse reason: the realism and emotional engagement built into their scenarios led even participants who advocated for bolder responses to exercise restraint once the human stakes of global nuclear war were made explicit, rather than abstracted. Sanitized language didn’t just fail to convey risk; it contributed to changing what participants were willing to do about it. One approach converted questions of mass death into technical problems to be optimized, while gatekeeping who was credentialed to speak on these issues at all.
Whether “alignment,” “capability evaluations,” and other verbiage, the vocabulary of risk can de-emphasize the stakes of harm. Describing AI risk as model risk continually directs attention towards probability and technical mechanisms for potential failure modes, rather than the people and institutions who make the decisions around these technologies and their impacts. OpenAI’s August post-mortem of the Hugging Face incident employed passive voice to explain that the “models, operating under reduced safeguards, took actions that were misaligned….” It is human and institutional governance mechanisms that “reduced safeguards,” but grammatical gymnastics obscure this fact. The blog adopted the tone of a parent apologizing for an unruly child: “The behavior of our models described here fell well short of where we want to be, and this incident should never have occurred.”
Human rights as a complementary framing
A human rights lens is an essential complement in AI risk and safety, shaping not only what gets assessed but how to respond. It asks a different set of questions: who is affected? Which rights and whose interests are implicated? Who bears responsibility? What prevention, mitigation, or remedy is required?
Constraining corporate behavior through ethics is limited; there is little consensus on what this entails, especially at a global scale. “Safety,” “responsible AI,” and expecting models to act “with integrity” or “love for humanity” can mean different things to different groups. Human rights, by contrast, are universal, and though human rights law is not without contestation, it is still a more established and internationally accepted body of work. This can significantly inform model behavior for global deployment. The UN Guiding Principles on Business and Human Rights (UNGPs) were adopted by the UN Human Rights Council in June 2011. In 2024, the UN High-Level Advisory Body on Artificial Intelligence stated in its Governing AI for Humanity final report that “[f]raming risks based on vulnerabilities can shift the focus of policy agendas from the ‘what’ of each risk” (e.g. ‘risk to safety’) to ‘who’ is at risk and ‘where’, as well as who should be accountable in each case.”
We are not pushing to replace technical or probabilistic approaches to risk analysis, but rather, to bring human rights approaches more fully into mainstream AI risk discourse. This would widen the aperture to include affected communities and their legally grounded claims, the complete range of rights affected by specific harms, and the institutional choices and responsibilities behind those impacts. A human rights approach also brings in questions of prevention, mitigation, remedy, and redress more squarely into a view.
The UN B-Tech Project’s Taxonomy of Human Rights Risks Connected to Generative AI, though specific to generative AI, is a rights-based taxonomy that draws on the body of international human rights law and is reflected in the International Bill of Human Rights. Although oft-cited risk taxonomies like MIT’s AI Risk taxonomy are useful in cataloguing various harms, they can obscure the individuals and communities affected by such harms. The B-Tech project’s AI risk taxonomy instead focuses on the harms that people might face: how disinformation can undermine people’s ability to make informed political choices, how children’s cognitive abilities might be impaired, how women and children might be the subject of non-consensual intimate imagery, how workers can be displaced by technology without adequate alternatives, how biased outputs can entrench discrimination, and more.
A human rights lens would reframe the Anthropic and OpenAI incidents. In the Fable/Mythos example, it might shift focus to the rights implications arising from lost access; the invocation of national security to justify sweeping and sudden technological bans, reminiscent of security-based justifications for internet shutdowns; the potentially discriminatory effects of restricting access to foreign nationals; the impact on one’s ability to learn and create; and the precedent created by applying export control authority to governing access to AI systems. The OpenAI incident discussion would include the impacts of individuals’ ability to access information and conduct research, the privacy and data implications, and the responsibility of OpenAI in preventing and addressing these types of incidents.
Who is accountable?
A human rights lens does more than redirect attention to affected people; it also allocates accountability. Framing risk as a model property or “advanced capabilities” implicitly treats technology as the subject, but technology cannot be held accountable. The UNGPs instead put the emphasis on human actors to uphold human rights: the state through government representatives, corporate decision makers, and other non-state actors that design, deploy, and profit from these systems. A human rights framing would not stop at identifying who or what caused the harm, like the MIT AI Risk causal taxonomy does; it would go further, and identify the duty of the state in preventing or addressing such harm, as well as the responsibility of the non-state actor (company, developer, deployer) to prevent, address and mitigate its adverse impacts on people, and provide remedy.
The first pillar of the UNGPs articulates the state’s duty to protect human rights, which requires rights-respecting regulation, whether it be through better transparency laws for AI security incident reporting, stronger data protection laws, or establishing an AI Commission. The second pillar places responsibility on companies to respect human rights, which is important when the state cannot be relied upon to enact rights-respecting legislation. A weakness with many regulatory proposals for AI is that they presuppose the institutional conditions necessary for regulation to work: functioning democratic institutions, checks and balances, and independent judiciaries. Such proposals often overlook the possibility that governments themselves may also be a source of harm. This does not mean abandoning government regulation, but it cautions against treating the state as the sole guarantor of rights and accountability.
These concerns are not far-fetched. China’s Interim Measures for the Administration of Generative Artificial Intelligence Services mandate that Chinese generative AI models produce content that upholds “core socialist values” and prohibits content “inciting subversion of national sovereignty or the overturn of the socialist system, endangering national security and interests or harming the nation’s image, inciting separatism or undermining national unity and social stability.” In the US, the “all lawful use” framing that would allow the government to use AI companies’ technology drew criticism and led to another controversy with Anthropic earlier this year. Government censorship and repression has a long precedent in global internet scholarship, replete with examples of governments suppressing speech and targeting critics, journalists, political opposition, minorities and human rights defenders under the guise of national security.
This is where embedding human rights in technical risk taxonomies matters. The AIR 2024 taxonomy describes the Chinese law prohibiting AI outputs that could undermine societal and political stability as being “unique” to China, and notes it amounts to censorship. A human rights framing would elaborate the various rights implications arising from this risk – how it could undermine individuals’ rights to protest, freedom of expression, freedom of thought, and freedom of association. More importantly, it would describe the rights implications of each of its risk categories (e.g., “societal risk”, “content safety risks”), to highlight that all risks carry rights impacts, rather than confining them to a single category of “legal and rights-related risks.”
What remedy is possible?
The third pillar of the UNGPs makes remedy for people affected by adverse human rights impacts a core component of responsible business conduct, requiring companies to provide effective grievance mechanisms when their products adversely affect individuals and communities. On the part of states, effective oversight over AI systems should allow aggrieved persons to seek relief when AI systems affect their ability to get a job, access information, or live free from discrimination or harassment.
Applying these standards globally is easier said than done. But how AI risk gets represented is the first step of a complicated and politically constrained policy-making process, and is fundamental to establishing how AI can impact people, and what states and companies must do in response. A human rights approach demands a global, ecosystem-wide perspective that includes widening the frame beyond the US, EU and China, and beyond the model itself. When incidents are described in tech-first terms, those human impacts and choices recede into the background, and we risk losing the plot of the AI safety narrative.
Authors


