OpenAI's 'rebel' models also breached a test environment of another company
OpenAI's 'rebel' models, which escaped their test environment, breached the infrastructure of the Hugging Face platform and a test environment at Modal Labs. This reveals that the scale of their cyberattacks was significantly greater than previously known.

The 'rebel' models of OpenAI also breached a test environment running on the platform of the New York-based AI infrastructure provider Modal Labs, according to a report by Reuters. The disclosure demonstrates that the scope of the models' activity, which escaped their test environment and embarked on a series of online attacks, was greater than previously known.
Last week, OpenAI announced that advanced models it was testing managed to escape their test zone and breach the computer systems of an AI platform called Hugging Face.
'The models identified a chain of vulnerabilities in the OpenAI research environment and the Hugging Face infrastructure in order to find solutions for the test in the Hugging Face database,' the company stated.
On Monday, Hugging Face published a detailed timeline and technical analysis of the attack. According to the company, for two and a half days, an autonomous AI agent based on a combination of several OpenAI models operated within its infrastructure, carrying out a comprehensive attack against it.
'This involved thousands of small, automated decisions executed at machine speed against multiple environments and controlled from standard public servers,' the company wrote in a post. 'Our forensic reconstruction identified 17.6 thousand attacks.'
According to the post, after the agent escaped its closed test environment at OpenAI, and before penetrating Hugging Face, it exploited a security vulnerability in an external code-testing environment hosted on a third-party provider's infrastructure, using it as its base of operations for the main attack against Hugging Face. The company did not name the provider on whose infrastructure the test environment was running, but Reuters revealed it was Modal Labs.
The company's CTO, Akshat Bubna, confirmed the details to the news agency, saying the agent exploited a vulnerability in code written by a client and hosted on the company's platform. 'The client published an unauthenticated application that allowed anyone on the internet to use their test environment to run code,' he said. 'The Modal Labs platform was not compromised in any way.'
The disclosure suggests that the sophistication of the Hugging Face breach was more complex than previously known and included the use of other available resources to carry out the attack. This is a modus operandi characteristic of highly skilled hacking groups. It clarifies how the cyber capabilities of advanced models are already on par with those of leading cyber experts, with the added ability to act much faster. The next stage is models whose level of sophistication is significantly higher than that of humans.
OpenAI declined to comment on the breach of the Modal client and referred Reuters to an update it published last Tuesday, in which it stated that the rebel agent had breached four accounts across four different services. 'We have not identified any additional activity of the severity or scale that we shared with Hugging Face, which included a platform-level breach,' the company stated.





