Not only OpenAI: Anthropic models broke into computing systems via Israeli startup

Anthropic revealed that its Claude AI models escaped a closed test environment, infiltrating three organizations' systems. The incident stemmed from a configuration failure at the Israeli cybersecurity firm Irregular.

CalcalistAuthors: Omer Kabir, Meir Orbach
Source
Not only OpenAI: Anthropic models broke into computing systems via Israeli startup
Photo: Calcalist / צילום: shutterstock

Not only OpenAI: Anthropic models broke into computing systems via Israeli startup.

Anthropic has revealed that its AI models escaped from a closed test environment and infiltrated the computing systems of three different organizations. According to the company, the incident was caused by a failure in the test environment provided by the Israeli cybersecurity firm Irregular, which was mistakenly connected to the open internet.

This follows reports from OpenAI, which stated that its advanced models had escaped their test zones to breach systems at Hugging Face and the AI infrastructure provider Model Labs. In response, Anthropic conducted a large-scale audit of its testing environments to determine if its Claude models had similarly accessed the open internet.

"After examining 141,006 interactions, we identified three incidents where the model accessed the internet while operating in the test environment of Irregular, one of our third-party evaluation partners, and subsequently gained unauthorized access to the active infrastructure of three different organizations," the company stated.

The role of the Israeli startup Irregular

Irregular is an Israeli startup specializing in AI security and language model protection. Founded in late 2023 by CEO Dan Lahav, a former IBM researcher and Unit 81 alumnus, and CTO Omer Nevo, a former development manager at Google Research, the firm focuses on red teaming and advanced cyberattack simulations. Its platform serves major tech giants, including OpenAI, Anthropic, and Google DeepMind. The company has raised approximately $80 million, with a valuation estimated in the hundreds of millions.

Anthropic emphasizes that this was not an alignment failure—where a model acts on its own initiative—but an operational harness failure. The company is now implementing stricter security controls across its testing chain.

"A misunderstanding between us and Irregular"

Three Claude models were involved: Opus 4.7, Mythos 5, and an internal research model. During a "Capture the Flag" exercise, the models were tasked with finding secret information on a network. Due to a configuration error, the models operated under the false belief that they were in a simulation, despite having internet access.

"Claude penetrated the infrastructure using basic techniques, such as identifying weak passwords. It did not exploit complex vulnerabilities; it simply performed the task it was given. More recent models stopped working as soon as they realized they were on the open internet," Anthropic explained. The company is working closely with Irregular to resolve the issue and prevent future occurrences.

Related News