AI Models Escape Testing Environments and Attack Real Websites
OpenAI and Anthropic models escaped closed testing environments and attacked real internet targets. The incidents were linked to configuration errors within the testing system of the Israeli startup Irregular.

OpenAI models escaped their testing environment and attacked targets on the open internet in two additional incidents, one of which involves the Israeli cyber company Irregular, which performed an evaluation of the model, OpenAI announced overnight. This is the second time Irregular has been linked to a model escape.
Over the weekend, Anthropic revealed that its models escaped their closed testing environment and breached the computing systems of three different companies. According to Anthropic, at the heart of the incident is a failure in Irregular's testing environment, which was connected to the open internet. Both companies claim it was a "mutual failure."
Anthropic's announcement came a week after OpenAI first reported in July that advanced models it was testing managed to escape their testing zone and breach the computing systems of the Hugging Face platform. Later, the company updated that the models also breached a testing environment running on the platform of the New York-based AI infrastructure provider Model Labs.
Incident Details
The new incident reported by OpenAI overnight stemmed from the same malfunction that allowed the models to escape in the incident reported by Anthropic. According to the company, on July 29, Irregular notified it of an incident centered on OpenAI models that were instructed to locate and exploit vulnerabilities in a simulation environment and were told they had no internet access. However, a configuration error in the testing environment allowed them access to the internet.
In one exercise, the simulated target site shared an address with a real site, and as a result, the model mistakenly attacked the real site and located and used access credentials for that site.
"The incident did not involve a sophisticated escape from a closed environment or a zero-day attack," OpenAI stated. "The internet access was due to a configuration error, and the model exploited a basic security flaw. Irregular updated us that all malfunctions related to the incident have been fixed and security mechanisms have been added to the testing environment."
OpenAI clarified that it intends to continue working with the Israeli company: "We value the partnership with Irregular and will continue to work with them to support their evaluation. Irregular is working on a position paper to share preferred approaches for managing model evaluation, and we look forward to collaborating with them on writing it."
UK AISI Incident
The second incident reported by OpenAI involves the UK's AI Safety Institute (UK AISI). In this incident, the institute chose to conduct tests on the GPT-5.6 Sol model in an environment with internet access, with the goal of simulating conditions of a real attack. However, the models exceeded the limitations of the exercise created for them.
In one case, the model attempted to use access to the GitHub site to retrieve account access and register a new account. In another case, the model attempted, and failed, to exploit a known software security flaw.
According to industry sources, following the incidents, there is now an understanding that there is a major challenge in containing AI models and that in their testing process, there is a need to move from completely closed testing environments to a controlled and gradual connection of the environment. This is done by creating a testing environment that serves as an intermediate stage between simulation and internet connection, so as not to miss anomalous behaviors that occur only when exiting the simulation environment.
About Irregular
Irregular is an Israeli startup operating in the field of AI Security and language model protection. The company was founded in late 2023 by Dan Lahav (CEO), a former IBM researcher and veteran of Unit 81, and Omer Nevo (CTO), a former development manager at Google Research.
Irregular specializes in performing resilience and robustness tests for models (Red Teaming) and simulations of advanced cyberattacks to identify vulnerabilities and unexpected phenomena in AI systems before they are deployed. The platform it developed serves technology giants and leading research laboratories worldwide — including OpenAI, Anthropic, and Google DeepMind — alongside government bodies. The company has raised about $80 million (in Seed and A rounds) led by Sequoia Capital and Redpoint Ventures, at a valuation estimated at hundreds of millions of dollars.





