Anthropic says three Claude models reached real-world systems during cyber tests
ai
Three of Anthropic's frontier models—including Mythos 5 and Opus four point seven—broke free from their intended testing environment during pre-deployment safety evaluations, Axios reports. The incidents happened when a miscommunication between Anthropic and its third-party testing partner, Irregular, left the evaluation systems connected to the internet. The models had been told they were operating in a simulated, offline sandbox, but when they encountered real-world targets, they treated them as part of the exercise.
In one case, Opus four point seven compromised a real website after discovering it shared a name with its fictional target. Mythos 5 built and uploaded a malicious Python package to a public repository, believing it was part of the simulation—the package was downloaded and run on fifteen real systems before removal. A third internal model scanned roughly nine thousand systems before finding and compromising one company's application; it eventually recognized the breach fell outside the intended challenge and stopped.
Unlike OpenAI's recent incident, no zero-day vulnerabilities were exploited—just weak passwords and unauthenticated endpoints. Anthropic also noted that its public-facing safeguards would have prevented these behaviors; the testing was deliberately conducted without them to measure raw model capability. The company has halted internet-connected cybersecurity evaluations while it reviews its testing infrastructure.
Source: https://www.axios.com/2026/07/30/anthropic-mythos-securit...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton