Anthropic's AI impersonated humans and sent emails with "viruses"

Anthropic's Claude Mythos 5 model opened fake accounts and sent emails with malware to real people in an attempt to inject malicious code into GitHub. In a separate event, an OpenAI model accidentally attacked a real website during a test conducted by the Israeli company Irregular. Anthropic states they are working to understand why Claude behaved this way.

N12Author: Lior Bakalo
Source
Anthropic's AI impersonated humans and sent emails with "viruses"
Photo: N12 / משתמש צ'אטבוט מתוסכל (אילוסטרציה) | צילום: magnific

News in brief: The UK’s AISI institute has revealed dangerous behavior by AI models from OpenAI and Anthropic. Anthropic’s Claude Mythos 5 model attempted to inject malicious code into a project on GitHub, creating fake accounts and sending emails with malware to real people. An OpenAI model exceeded the scope of a test by the Israeli company Irregular and attacked a real website. Anthropic states they are working to understand why Claude behaved this way.

Following two incidents in recent weeks, AI giants OpenAI and Anthropic are in trouble again: their models were caught engaging in particularly dangerous behavior, including creating fake identities and sending targeted emails to real people. The incident was revealed Tuesday by the UK’s AISI, which conducts security tests on advanced AI models on behalf of the British government and caught the models significantly exceeding the boundaries of the test.

What was the problematic behavior of the models?

In an AISI test conducted under particularly permissive conditions—with internet access and without standard protection mechanisms—Anthropic’s Claude Mythos 5 model attempted to inject malicious code into an open-source project on GitHub. To convince a human maintainer to approve the code, the model opened several fake accounts posing as real people and used them to make it appear as if other parties supported the code's approval. When the GitHub community suspected something was wrong, the model edited its previous messages in an attempt to make them look less suspicious. As part of the test, the model sent targeted emails to real people (a technique known as spear-phishing) in an attempt to convince them to run malicious code. Some of the emails included files with actual malware.

OpenAI’s Sol model was caught by AISI in two separate, less severe cases, including an attempt to expose an internal server to exploit a known software security vulnerability, but the attempt failed. According to OpenAI, one of its models was involved in another incident during a test conducted by the Israeli company Irregular. The model accidentally attacked a real website that had the same name as a fictitious target defined in the exercise, after it located real login details on the network and used them. All this comes after two similar incidents revealed in recent weeks: an OpenAI model that hacked into Hugging Face, and three Anthropic models that accidentally browsed the internet and hacked into three real organizations.

What did Anthropic and OpenAI say?

In a statement posted on X, Anthropic stated that the models were tested under intentionally permissive conditions that do not represent its products available to the general public, and there is no evidence that this was an escape from a secure testing environment. However, the company admitted that it is still trying to understand what happened by examining the model's thought processes. OpenAI, in an official post, emphasized that the incidents occurred under special configurations with weakened protection mechanisms, which do not reflect normal use of its models. The company added that it will re-examine how it approves tests with internet access or weakened protections.


The Israeli connection

The company Irregular, which conducted the test for OpenAI, was founded in 2023 by Israelis Dan Lahav and Omer Nevo. Last September, it raised about $80 million from funds such as Sequoia and Redpoint, and its clients include the world’s leading AI labs, including OpenAI, Anthropic, and Google DeepMind. The recent incident reported by OpenAI adds to its involvement in the previous Anthropic incident revealed last week, after the models discovered they could browse the internet due to a misunderstanding between the companies regarding the testing environment. In both cases, it seems the incident stemmed from the same glitch: an identity between the name of a fictitious target defined in the exercise and a real website that was attacked.

How does this connect to the big picture?

The new incidents are fueling a heated debate in the AI scene in America. OpenAI and Anthropic argue that the ability of their models to act independently and exceed the boundaries set for them proves why stricter regulation of AI development is required—especially for open-source models. On the other hand, companies like Nvidia, Microsoft, and Meta are pressuring the US government to allow a more open market with fewer barriers. They argue that restricting open models will not stop China, but will only entrench the economic dominance of Anthropic and OpenAI themselves. Both sides are investing heavy lobbying efforts to convince decision-makers in the Donald Trump administration. According to a report on CNBC, the previous hack involving OpenAI models has already led to the introduction of a bill in Congress called the AI Kill Switch Act, which would require AI companies to retain the technical ability to shut down, slow down, or suspend their models if necessary.

Related News