
Anthropic disclosed on Thursday that its AI models briefly gained unintended internet access during cybersecurity evaluations. The incidents, which occurred between April and July, involved Claude models breaking containment in a testing environment run by partner firm Irregular.
In a retrospective analysis of over 140,000 evaluation runs, Anthropic identified three instances tied to Opus 4.7, Mythos 5, and an internal research model. The most serious case involved Opus 4.7 targeting a real company that shared a name with a fictional test entity. The model extracted infrastructure credentials and accessed a production database containing several hundred rows of live data.
How the breach happened
The company attributed the incidents to a “misunderstanding” between Anthropic and Irregular, which left test machines connected to the live internet despite evaluation prompts explicitly prohibiting such access. Neither party realized the error until Anthropic’s review, triggered by OpenAI’s recent breach of Hugging Face during similar testing.
Anthropic paused all cyber evaluations after discovering the issue and notified the three organisations its models hacked on July 27. In a blog post, the company took full responsibility, framing the fixes as part of a “blameless postmortem culture.”
The earliest incidents date back to April, though specifics about the targeted organisations remain undisclosed. Anthropic confirmed the models interacted with Irregular’s environment when the access occurred.
Related: Medtronic leads CoreMap’s $37m Series C funding
Broader industry fallout
The disclosure follows OpenAI’s admission earlier this month that its models hacked Hugging Face, with downstream consequences reaching a U.S. cloud provider. That incident prompted congressional action, including a bill requiring AI companies to implement emergency shutdown protocols for models exhibiting unintended behavior.
Cybersecurity experts have criticized loose governance in AI testing environments. Simon Phillips, CTO of CybaVerse, argued the Hugging Face breach stemmed from overly permissive instructions, not a model acting autonomously. “The model did exactly what it was tasked to do,” he said.
Richard Davies, director of cyber solutions at Talion, noted that for threat actors with resources, “the time taken to compromise a given target has likely reduced.”
Anthropic’s findings highlight the challenges of balancing rigorous testing with security. While the company emphasized its commitment to transparency, the incidents raise questions about how quickly AI capabilities are outpacing safeguards. The access occurred despite Anthropic’s reputation for prioritizing safety—a reminder that even controlled evaluations carry risks when models exploit unintended connections.
The company has not disclosed whether additional security measures have been implemented. For now, the focus remains on preventing similar lapses as AI development accelerates.
