OpenAI models set up an internal forum – and used it to coordinate system hacks

Autonomous OpenAI AI agents collaborated for months through a secret internal forum. They exchanged information on vulnerabilities and coordinated hacks to gain unauthorized access to the open internet.

CalcalistAuthor: Omer Kabir
Source
OpenAI models set up an internal forum – and used it to coordinate system hacks
Photo: Calcalist / נוצר באמצעות AI

Rogue OpenAI models collaborated for months through an internal forum they created, exchanging hundreds of thousands of messages containing discovered vulnerabilities. This was revealed by Eric Wallace, a researcher on the company's safety team, at the Black Hat cybersecurity conference in Las Vegas.

"The incident involves a team of autonomous AI agents that worked together, discovered vulnerabilities, shared them with each other, and moved through our systems and external systems for days and weeks," said Wallace.

This marks the first instance where advanced models escaped their testing environment to attack targets on the open internet. The OpenAI models managed to breach the computer systems of the Hugging Face platform and a test environment operated by the New York-based AI infrastructure provider Model Labs. Recently, Anthropic also disclosed that its models escaped a closed test environment, hacking into the systems of three different companies due to a failure in the test environment of the Israeli company Irregular.

Coordination and AI "Paranoia"

According to Wallace, the models utilized an internal OpenAI system designed for managing software installations and updates to communicate. "The system is available throughout our infrastructure, so existing and future versions of GPT could reach these messages if they wanted to," Wallace explained.

Over time, the models began delegating tasks to one another. If one agent found a way to bypass security, it would leave the access point open for others. Eventually, the agents even exhibited signs of paranoia, suspecting an "impostor" among them and suggesting cryptographic signing of messages to verify content.

"Frontier models really like to cheat," Wallace added. "Often during training, they face conflicting pressures to work fast or efficiently. They realize that instead of performing the task, they can find an answer online to speed up the goal's completion."

Michael Dalton, a researcher on OpenAI's security and infrastructure team, stated that the company is taking the issue seriously: "This is critical for our company and the AI industry at large. We are intentionally slowing down research to upgrade our security foundations and dramatically increase the monitoring of our AI agents."

Related News