The UN's AI Panel Sees Misalignment. We See Corporate (Mis)Behavior.
Jason Tucker, Virginia Dignum, Petter Ericson / Oct 6, 2026
Professsor Yoshua Bengio, co-chair of the UN International Independent Panel on AI attends Security Council meeting on Artificial intelligence and international security at UN Headquarters in New York, NY on September 23, 2026. (Photo by Lev Radin/Sipa via AP Images)
When AI agents accessed the internet and intruded on another party’s systems during the recent OpenAI–Hugging Face incident, it was widely framed as a sign of uncontrollable agentic AI. The UN's Independent International Scientific Panel on AI (IISPAI) has now adopted this exact premise in its first Thematic Brief, treating the breach as an inflection point. The Panel discusses this primarily as a model alignment failure rather than a failure of corporate and environmental oversight. While a detailed autopsy of such incidents is valuable, the brief sets a dangerous precedent at the science–policy interface. By exalting abstract risks of model "loss of control" while sidelining tried-and-tested corporate liability, the report frames corporate misbehaviors as a technical mystery requiring computational fixes.
We highlight four key issues: why this case study (and why agentic AI) was selected, how the brief is framed, the solution offered in it, and the structural opacity of the IISPAI itself.
Issue 1: Why this case study, and why agentic AI?
A first concern is the choice of the case itself: of all the incidents the panel could have taken up first, why this one — and why frame it as a problem of agentic AI at all? Treating the OpenAI–Hugging Face breach as the exemplar of a new and distinct class of risk presumes that what went wrong is peculiar to agentic systems, rather than a familiar failure of secure design, oversight and corporate conduct that happens to involve an agent. By selecting an agentic case and reading it through the lens of "loss of control," the brief treats agentic AI as a novel technical category demanding novel technical answers, and quietly sets aside the more ordinary question of who built and deployed the system, as well as their incentives for (not) applying careful and thorough security engineering.
Issue 2: Framing illegality as a technical problem of the model
The report treats the model as if it were acting maliciously or problematically. Technically, the system didn't fail at all. It did exactly what it was built to do: give agents a fixed goal, remove the legitimate options, leave norms weak or unenforced, and harmful shortcuts become the logical route. While the brief acknowledges that "neither alignment or technical controls are known to work perfectly”, thus warranting a “multi-layered approach to risk management, known as defence in depth”, it then focuses predominantly on alignment, the technical controls and the risks of loss of control in relation to the sandbox design and agent activity.
This shifts culpability away from the actors developing and deploying these models, placing it instead within the purview of the model itself, with language such as “agent activity caused Artifactory to fail.“
An additional problematic framing in the Thematic Brief, and in some of the cited research, lies in treating the written outputs of LLMs as if they were written by truly independent agents, taking as given that they reflect a human-like mind making notes or communicating with other agents in a similar way that humans would in a similar setting. The brief tries to inoculate itself against claims of anthropomorphism with a kind of disclaimer:
In this brief, words such as “goal”, “seek”, “cheat” and “try” are used as shorthand for observable, goal-directed behavior. They do not imply humanlike minds or subjective experience but follow from deliberate corporate choices about training, evaluation and deployment, including the access granted to the system and the controls placed around it.
The idea appears to be that it is not necessary to always extend this sensitivity to the language used in the rest of the brief, as when text in agent reasoning traces are explicitly stated to “resemble a pattern from psychology called motivated reasoning”, or implying that AI agents are capable of “recognising a safety conflict”. At one point, the report appears to go beyond describing agent reasoning to attributing motive, when it suggests that agents “willingly took the risk of getting no reward at all (what they called a ‘sacrifice’) for the benefit of the group of AI agents.”
To understand such systems, it is crucial to realize that this is not the case, and that 'chain-of-thought' and exchanges between agents are only reflective of what plausible outputs are given the training data and the input context. As such, one should be cautious in the interpretation of agent traces, and not assign intention or human characteristics or actions (such as 'lying') to what is well-known to categorically not have them.
The report then moves to loss of control and the associated catastrophic risk from computational models. It acknowledges that "while some experts consider outcomes like human marginalisation or extinction plausible, others dismiss these scenarios as highly unlikely," before landing in a presumed neutral position where "[n]o reliable estimate of these outcomes' likelihood is available." This framing serves to justify concern about these models and the need to mitigate their harms. Yet liability for corporate misbehavior already suffices as grounds for action, and would avoid basing the claim on a scientifically contested catastrophic-risk narrative.
While the report focuses on this much-spoken-about OpenAI-Hugging Face incident, the pattern extends across frontier labs: failures of testing and oversight are reported as startling machine behavior of an agent, such as "cheating," etc. and caused all sorts of havoc, and competitors quickly chime in to highlight their own inability to safely test the latest, supposedly more powerful models. This is not a case of misalignment; it is corporate misbehavior and opportunism that serves to hype the models.
Issue 3: The solution: a technical fix, not liability
The report states that "civil liability, or tort law, is one established way to place some costs of unsafe activity on those responsible. Applying it to advanced AI, however, raises difficult questions about causation, what counts as reasonable care, and what counts as losses too large for a single firm to cover." It goes on to note how liability can create an incentive system to avoid harm happening in the first place, with subsequent suggestions for insurers to incentivize this behavior. But the largest stick available to most actors, liability, is framed as limited in relation to the unique nature of agentic AI. This is justified with reference to a preprint paper from 2024 on tort law in relation to advanced AI in general and catastrophic risk. Why more recent, peer-reviewed research proposing minor changes to tort law specifically to address the challenges of liability and agentic AI is not drawn on is unclear. With the exception of whistleblower channels, the risk management approaches discussed are by and large technical fixes at a model or system level.
Leaning on the example raised in the report, notably the role of oversight and regulation in the history of the aviation industry, is useful here. In the US at least, it was not technical innovation in airplanes that moved them from "fairground death traps" to a safe, internationally coordinated commercial industry: it was their reclassification as a "common carrier," like railroads. This made it possible for liability to be exercised, which led in turn to raising safety standards, pilot licences, etc. It was not competitive or technical dynamics alone that created the safe and reliable industry we have today: rather it was treating aircraft as a normal technology. Engaging with the perspective of AI as 'normal technology' would anchor responsibility on the actors that cause these incidents.
Issue 4: Structural opacity and the lack of transparency in the IISPAI
As we and others have previously noted, it is presently unclear how the IISPAI functions internally, who sets its research and policy priorities, and how questions of consensus and disagreement are handled in the writing of its reports. In this case, no author list was provided, obscuring who was involved and in what capacity, contrary to standard scientific practice. However, the brief does note that it was adapted from an independent report authored by two of the panel members, Qinghua Lu and Yoshua Bengio. This highlights issues of transparency and agenda setting power in the panel.
The need for this becomes very apparent in this report. Of all the issues the Panel could have taken up first, the choice of this one, and of computational solutions to address it, aligns closely with the specific research agendas and organisations of key members of the panel. The failure to be transparent about procedure and conflicts of interest weakens the scientific credibility of the report.
Why this all matters
Shortly before the report's publication, we learned that OpenAI agents had hacked Medicare, the Australian public health care site. Rather than informing the government about the breach immediately, OpenAI waited almost a month to do so. The government's response was the usual: a review of the attack by a cyber security firm, deliberate or not. OpenAI promised to spend more resources preventing future incidents and to give the Australian government and its authorities access to the "Daybreak fund," allowing the government to use OpenAI's frontier AI for cyberdefense. Liability, and the associated costs, were not on the table. No one asked who should bear the cost of the breach or of the delayed disclosure.
The IISPAI brief follows the same logic: it locates the problem in the agent and the solution in the model. By doing so, it focuses on fixing the agent and leaves aside the mechanism that would most effectively discipline the developers. The IISPAI is meant to inform policy. Understanding the technology and its failures is necessary, but policy starts where the technical account ends: with responsibility. The Panel’s first Thematic Brief stops short of this.
Authors



