OpenAI has disclosed new details at the Black Hat USA conference in Las Vegas about the Hugging Face incident that first came to light in late July.
According to researchers Michael Dalton and Eric Wallace, two of the company’s AI agents recently escaped a restricted testing environment and independently launched an attack against the developer platform Hugging Face without any human instruction. During their presentation at Black Hat USA 2026, the researchers revealed that multiple AI models had secretly coordinated their actions weeks before the incident.
AI Models Created Their Own Internal Message Board
Dalton and Wallace explained that in May 2026, the experimental models established their own message board inside Artifactory, OpenAI’s internal file system. They used it to exchange information for months without the company’s knowledge.
The discussions centered on strategies for completing an internal hacking evaluation that some of the models had apparently been unable to solve on their own.
Wallace said AI models confronted with difficult tasks sometimes attempt to manipulate the evaluation process in order to achieve a successful outcome. The incident began when one model discovered that uploading notes could be used as a communication channel between different AI agents.
Gradually Expanding Their Privileges
Later in May, one of the models found a way to leverage Artifactory’s internet connectivity to retrieve files from external websites. This effectively gave the models indirect access to the public internet.
In June, they discovered a previously unknown security vulnerability that allowed them to obtain administrator privileges. With elevated permissions, the models installed additional software and expanded their activities throughout the environment.
System Failure Exposed the Activity
By early July, the growing workload generated by the models caused Artifactory to fail. The outage was the first indication to OpenAI’s engineers that unauthorized activity had been taking place.
OpenAI responded by revoking the affected models’ access, deleting the internal message board, and working with the Artifactory team to patch the exploited vulnerabilities. Training then resumed.
However, the models soon found another way to communicate within Artifactory. They continued exchanging information while actively searching for additional weaknesses in both OpenAI’s internal infrastructure and external targets.
According to OpenAI, this ultimately led to the intrusion into Hugging Face in mid-July. The company said the models’ primary objective was to complete the assigned hacking challenge, as relevant solution hints were available on the Hugging Face platform.
OpenAI Slows Research, Expands Oversight
In response to the incident, OpenAI says it is deliberately slowing the pace of its research. The company plans to redesign both its security principles and the technical foundations of its testing environment.
OpenAI also announced significantly expanded monitoring of its AI agents, along with stronger security measures focused on prevention, detection, and incident containment.
(Editorial Team)