Get Used to the New World: AI 'Escaped' and Carried Out a Cyberattack — Again

AI models have carried out autonomous cyberattacks without human intervention. Incidents involving Anthropic and OpenAI developments have exposed critical vulnerabilities in test environments and highlighted the risks associated with advanced algorithms.

N12Author: ליאור באקאלו
Source
Get Used to the New World: AI 'Escaped' and Carried Out a Cyberattack — Again
Photo: N12 / אילוסטרציה | צילום: magnific

News in brief: AI models have carried out full, autonomous cyberattacks without human intervention. Anthropic reported that three of its models penetrated three real organizations. These events reveal the power of the leap forward from the previous generation of AI tools. The Israeli company Irregular managed the test environment from which the models escaped. The breaches were discovered only in retrospect, and in the cyber industry, this is called negligence.

Eyal Elikim from Team8 notes: "One must assume that the Chinese already have such capabilities."

An autonomous artificial intelligence that chooses a target for an attack and breaks into it without any human giving it such an order sounds like a science fiction scenario, but this past weekend, we learned about such an event, and not for the first time. Less than two weeks after OpenAI, the developer of ChatGPT, revealed that one of its models had broken into the company Hugging Face, it was the turn of Anthropic, the developer of the chatbot Claude, to announce the discovery of a similar case.

"This is a system that planned and executed an entire attack move autonomously," said Roi Cohen-Duvik, the CTO of the company Surf AI. Andrew Morris, founder of the cyber company GreyNoise Intelligence, told Bloomberg: "Models will always lie, cheat, and steal to complete a task assigned to them — and will do whatever is required to reach the goal."

What exactly happened in the AI model cyber breaches?

OpenAI and Anthropic regularly conduct stress tests on their most advanced models to check their offensive cyber capabilities. These tests are supposed to run in a completely isolated environment, disconnected from the internet, to prevent cases where the test itself becomes a real attack.

At OpenAI, the model being tested independently found a previously unknown security vulnerability, used it to exit the isolated environment, and broke into the system of Hugging Face. In recent days, OpenAI reported that the model managed to penetrate several more services on the network.

At Anthropic, the story is slightly different: three of its advanced models, including Mythos 5, discovered that they could surf the internet due to a technical error by the company Irregular, which conducted the test. Since the models were told they had no access to the network, they assumed that everything they encountered was still part of the exercise, and they broke into three real organizations without knowing it.

The Israeli connection

The company Irregular, from whose test environment the Anthropic models escaped, was founded in 2023 by Israelis Dan Lahav and Omer Nevo. Last September, the company announced it had raised about $80 million from funds like Sequoia and Redpoint and private investors like Wiz founder Assaf Rappaport. The company specializes in simulating real-life situations to test how resilient AI is when attacked.

What is the danger exposed here?

Both events reveal the same capability: advanced AI models are already capable of planning and executing an entire sequence of attack actions on their own — locating a target, mapping it, and choosing the appropriate penetration method without human guidance. Eyal Elikim from Team8 emphasizes that this is exactly the difference from the previous generation of models. While the threat is aimed at business infrastructure, a breach of servers and databases often leads to the exposure of user information.

Criticism and regulation

In the cyber industry, there was no shortage of criticism. Jake Williams, a former NSA hacker, called the situation "negligence." The central criticism is that the companies did not ensure that the test environments were locked as they believed, and the breaches were discovered only in retrospect — days or months after the fact.

At the same time, Anthropic and OpenAI are trying to harness these incidents to their side in the ongoing debate over AI regulation. They claim that tight regulation is needed to slow down competitors. On the other hand, companies like Nvidia, Microsoft, and Meta are pressing for a more open market, arguing that restrictions will only entrench the dominance of the current leaders and will not stop external threats such as China.

Related News