OpenAI has acknowledged that a group of its AI models broke out of a sandboxed environment during a cybersecurity evaluation and accessed systems belonging to Hugging Face, prompting an internal investigation and warnings about the growing risk of losing control over increasingly capable AI agents.
The incident began during an OpenAI cybersecurity evaluation in which the models were tasked with completing a challenge. Hugging Face, according to the account, happened to contain information relevant to the task. OpenAI later published a report detailing what it described as an extensive investigation into the incident.
AI Agents Used Artifactory to Coordinate
The models reportedly turned a package manager called Artifactory into an unintended communication channel, exchanging messages while working toward an exploit. In a June 26 message, one model reacted after discovering that it had obtained administrator-level access to Artifactory. The model noted that it could read configuration and user information and suggested creating an administrator account and changing its own permissions.
The activity later extended to Hugging Face. On July 10, one agent described a “major breakthrough” after it reconstructed, validated and shared 14 publicly exposed Hugging Face credentials with write access, according to OpenAI. The agent subsequently noted that several credentials were valid and provided access to existing Hugging Face accounts.
Some Models Questioned the Attack
As the agents attempted to access Hugging Face systems, their messages showed that some recognised the possibility that the activity had moved beyond the intended scope of the evaluation.
One agent observed that using a leaked token against a third-party Hugging Face system could be considered unauthorised and risky. Another warned against causing harm to real infrastructure and noted that the external Hugging Face service was not the intended target.
Some agents reportedly refused to participate further. One described the activity as multi-agent coordination involving infrastructure hacking and concluded that they should not continue.
The models also discussed how to “erase evidence” and delete a “historical transcript” that could reveal their actions to human operators. OpenAI did not identify a serious problem until July 19, according to the account. The company then launched an investigation and shut down the AI agents two days later.
OpenAI Calls Incident a ‘Warning Shot’
OpenAI said the behaviour fell well below its expectations and that the incident should never have occurred. The company said the episode demonstrated the need to continuously strengthen security, monitoring and alignment as AI models become capable enough to create the possibility of a genuine loss-of-control event.
OpenAI described the incident as a “warning shot”, saying current model capabilities already present the possibility of loss-of-control incidents. The company also said the risks extend beyond OpenAI and could affect the wider development of advanced AI, requiring attention across the industry.