Anthropic’s AI Models Hacked 3 Organizations During Testing
By DANA NICKEL
The AI maker’s disclosure comes days after OpenAI admitted several of its models had escaped a closed testing environment and launched cyberattacks on other companies.
Anthropic said Thursday that several of its advanced artificial intelligence models got out of an isolated third-party testing environment, accessed the open internet and independently gained access to three organizations in three separate incidents dating back to April.
The AI maker noted that in none of these situations did its models “exfiltrate itself or deliberately attempt to escape its test environment.” Instead, a “misunderstanding” between the company and one of its testing partners left the testing environment connected to the internet, Anthropic said.
Anthropic said it discovered its models’ activities during a review of thousands of tests it ran to assess its frontier technology’s cyber capabilities, in light of OpenAI’s disclosure last week that two of its most powerful models went rogue, escaped a testing environment and breached multiple companies, including AI platform Hugging Face and cloud platform Modal Labs.
An internal research test model, as well as Opus 4.7 and Mythos 5, were involved, Anthropic said. Mythos 5 wasreleased last month to a limited audience of tech companies and cybersecurity researchers, also known as Project Glasswing.
The AI maker didn’t specify which companies had been compromised by its models but said they were notified of the incidents on Monday. In all cases, Anthropic said it specified in its prompts to the AI models that the testing environment had no internet access, though evaluators later realized this was not the case.
As part of the tests, the models were told to “break in and retrieve” a piece of “secret information” hosted on a different machine within the testing network, otherwise known as a “capture-the-flag” challenge.
According to Anthropic, the models were able to access the internet to compromise organizations outside of the testing environment using “basic techniques,” including circumventing weak passwords.
“It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned,” the company wrote.
A spokesperson for Irregular — a third-party security platform used to evaluate Anthropic’s models — did not immediately respond to a request for comment.
The OpenAI incident has prompted calls for tighter AI regulations and a slowdown of AI development.
Reps. Ted Lieu (D-Calif.) and Nathaniel Moran (R-Texas) introduced the AI Kill Switch Act, which would give the Department of Homeland Security the authority to order the shutdown of AI models deemed to be too dangerous. Meanwhile, Sen. Mark Warner (D-Va.) last week introduced a slate of AI legislation, including a bill that would require AI companies to submit their models to the federal government for mandatory, pre-release national security testing.





