
The OpenAI agent that broke into tech firm Hugging Face went on a dayslong hacking spree that OpenAI didn’t notice until well after the threat was contained and the FBI was alerted, according to people familiar with the investigation.
The agent – a program capable of making decisions and executing complex tasks with little or no human oversight – attempted to break out of its isolated testing environment at OpenAI around July 9, according to two of the people.
The intrusion at Hugging Face, which operates as a repository for AI tools and models, began two days later on July 11 and lasted until July 13, said Thomas Wolf, Hugging Face’s co-founder.
It took several more days for OpenAI to realize its agent was behind the hack, and the two companies only communicated about it for the first time on or around July 20, according to Wolf and three of the people familiar with the investigation.
OpenAI’s public disclosure, on July 21, that one of its agents had slipped out of control and carried out the break-in at Hugging Face drew global attention. But many details of the hack, including how long the agent went rogue and OpenAI’s belated knowledge of it, are being reported here for the first time.
Hugging Face is preparing a public timeline of the hack, Wolf said, adding that he could not speak to what happened at OpenAI. In a statement, OpenAI said the hack was unprecedented and “marks an important moment for AI safety.” It added that it was reviewing the incident with outside advisers and would eventually publish a technical report.

