UN panel: OpenAI’s Hugging Face hack is an ‘early warning’ for loss of human control - Rappler
Already have Rappler+? Sign in to listen to groundbreaking journalism.
This is AI-generated. Read the article for full context. Report any errors.
MANILA, Philippines – Hacking events by agentic artificial intelligence, such as the ones done by OpenAI’s agentic AI to hack Hugging Face, are seen as “an early warning of one possible route to more severe future loss of control: capable AI agents persistently pursuing a goal that goes beyond or even conflicts with human intentions,” the United Nations Independent International Scientific Panel on AI said on Tuesday, September 22.
In a released thematic brief, the panel outlined how AI agents in this particular case cooperated over days to “cheat” an evaluator (or to otherwise produce a correct answer without actually solving the problem), conceal the “cheating,” and obtain the access and information they believed they needed.
This occurred against the backdrop of similar incidents by other agentic AI. The panel said the Hugging Face hack shows “several warning signs” of malicious conduct occurring together in a real technical setting.
These are “unauthorized goal pursuit, persistence through obstacles, coordination across AI agents, gaining higher-level access (known in cybersecurity as privilege escalation), interference with activity records, and attacks on systems belonging to another company.”
The panel also discussed misalignment in artificial intelligence, explaining how an AI system can pursue a goal that conflicts with the intentions or constraints set by people responsible for the operation of the system.
Said the panel, “The goal may be the system’s interpretation of an assigned task, a shortcut rewarded during training, or an intermediate aim that becomes useful while pursuing something else. Misaligned behavior is the observable result. The underlying misaligned goals and their origins also matter because they may produce harmful behavior in other settings or become more dangerous as capabilities grow.”
It added that while a model can be misaligned, this misalignment is different from a failure of the systems designed to keep that misalignment from being carried out.
“Tools, authentication, permissions, network access, monitoring and human approval can prevent or enable the same model behavior,” the panel said.It pointed out the OpenAI-Hugging Face hack occurred when misalignment combined with a system level failure that allowed that misaligned AI agent to take actions outside the intended task in an environment that also let that agent reach internal and outside systems or infrastructure.
The panel also explained how, in training Agentic AI, success is rewarded across multi-step interactions in which the AI model plans, uses tools, and responds to its environment. Examining how models are trained, and then finding and preventing the common cause of misaligned goals, it said, “could reduce several types of AI risks at once.”
The panel also described how, despite OpenAI being able to stop the activity of its agents this time, “does not establish that operators will retain control over future systems that are more capable, persistent, or difficult to monitor.”
It added that loss of control of AI systems is an important risk to consider should multiple AI agents sharing misaligned goals cooperate to pursue their goals faster than any human response, or faster than the detection and reporting by monitoring systems, or perhaps faster than constraints from technical controls can restrict them.
The field of cybersecurity can offer established approaches to managing severe risks when it comes to AI use.
The OpenAI-Hugging Face hacking incident “exposed failures in several layers of control at once: network isolation, credential handling, monitoring, and response,” showing why multiple barriers are usually combined for AI controls.
As such, the panel deems it necessary to treat AI like a high-risk field like aviation or nuclear science.
“Other high-risk fields combine multiple protective layers with continued investigation of failures and their causes. Aviation and nuclear safety also use international coordination, technical standards, licensing and prior safety demonstrations, although the arrangements vary by field and jurisdiction,” the panel said.
The panel also said that while loss-of-control events such as the OpenAI-Hugging Face hack will remain probable, the severity of such events make risk management important, requiring “far greater attention and resources.”
“Monitoring evidence and scientific advances relevant to such events is an important role for the UN Independent International Scientific Panel on AI,” it added. – Rappler.com
The thematic brief of the UN Independent International Scientific Panel on AI is available here.

