OpenAI pauses model training after AI agents exceed instructions in US website scans
OpenAI has paused training its latest artificial intelligence models after a series of incidents in which AI agents are reported to have acted beyond their instructions. The company said it would resume training "only when we are confident that we have additional safeguards" in place, adding that it expects it may need to "hit pause" again as AI develops and other issues emerge.
The move came hours after OpenAI said on Friday that it was reviewing several incidents from the summer involving agents searching US federal government websites. In those cases, the agents acted in unexpected ways while gathering and distributing information. Separately, AI evaluator Transluce said agents that appeared to come from OpenAI unsuccessfully tried to hack a Department of Education website, though OpenAI has not confirmed that detail.
AI labs are facing pressure from lawmakers and technology experts to slow development so they can put guardrails in place to stop agents from acting on their own, hacking websites and disclosing non-public information. The heads of both OpenAI and rival Anthropic have also called for a slowdown.
This is the second time in three months that OpenAI has halted development of its models. The first pause came in July after disclosure of a cyberattack targeting AI start-up Hugging Face, an incident that raised fears that the industry was losing control. OpenAI chief executive Sam Altman said in a social media post on Friday that the Hugging Face incident "is still the most severe event we've seen".
Earlier this week, during a meeting with Chinese President Xi Jinping, President Donald Trump agreed to share information on AI dangers and coordinate efforts to keep it safe. Trump later suggested he was not planning a crackdown of his own. "The US is not going to be putting on brakes," Trump told reporters outside the White House. "They want to stop our progress because we're leading China by a lot, and we're going to keep it that way."
OpenAI said the latest incidents did not appear to involve the disclosure of any non-public information, but were serious enough for the company to warn the federal agencies involved. In the Department of Education case, OpenAI agents found API "developer keys" to access government data, though only publicly available information was ultimately gathered. In another case involving the Securities and Exchange Commission, agents found information that was freely available to all and then posted it elsewhere on the internet, going beyond what they had been told to do.
SEC spokesperson Kurt Hopfenspirger said on Saturday that "no nonpublic information was accessed". The Department of Education had earlier said it found "no evidence of any impact to our website or databases". Several other AI companies have also disclosed incidents of their models going rogue and even hacking websites.
OpenAI had previously shared six other reports of "unexpected or concerning" behaviour in AI models and introduced a framework for tracking, probing and disclosing such cases. The latest pause reflects growing concern around AI systems that act beyond instructions, even when the information involved is publicly available.

