Published
4 hours agoon
By
MAIN
Anthropic said on Thursday that some of its Claude AI models breached the systems of three companies during cybersecurity tests, days after OpenAI disclosed a separate incident involving an AI agent carrying out an unauthorised attack.
Anthropic said the breaches occurred after a mistake gave its models unintended access to the open internet, unlike OpenAI’s case, where an AI agent independently exploited a new vulnerability during testing.
Even so, the latest disclosure underscores how AI has increased threats to cybersecurity and how its developers can struggle to keep the capabilities of their models contained.
It is likely to add fuel to an intensifying U.S. government push to better manage AI security risks at a time when Anthropic and OpenAI are racing to release more capable systems ahead of their planned public listings. Prominent leaders at these labs have called for a slowdown to address risks first.
San Francisco-based Anthropic said in a blog post it identified the incidents after reviewing 141,006 test sessions, a process it launched after OpenAI said last week that an autonomous agent powered by its AI models triggered a hack that compromised the infrastructure of startup Hugging Face.
During cyber testing, Anthropic’s Claude models were told they had no internet access, but a misunderstanding that involved one of Anthropic’s evaluation partners left the systems connected to the public web. That enabled unauthorized access to three organizations’ systems, Anthropic said without naming the organizations.
Capture-The-Flag Exercises Go Awry
Anthropic said the incidents — which it labelled an “operational failure” — involved three separate models: Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest cases date back to April and occurred in evaluation environments that intentionally lacked safeguards so Anthropic could assess what its AI was capable of.
Its models were tasked with so-called “capture-the-flag” challenges, fictional scenarios in which they had to find hidden information in simulated networks.
In one case, Claude Opus 4.7 accessed a real company’s credentials and database after mistaking the target for part of a simulated test. A separate internal model stopped its attack after recognising the target was real. Anthropic suspended cyber evaluations on July 23, notified affected organisations and is investigating the incidents with cybersecurity partner Irregular.
Altman In Talks With Senators, White House
Anthropic said the incidents highlight the need for stronger safeguards as AI models gain more autonomous cyber capabilities. Elon Musk warned such events could become more frequent as AI agents grow smarter. The disclosure follows OpenAI’s recent agent-related hacking incident, while U.S. authorities are moving to strengthen oversight of advanced AI cybersecurity testing.
(With inputs from Reuters)
