Anthropic has disclosed that multiple versions of its Claude artificial intelligence models inadvertently breached the systems of three separate organizations during cybersecurity testing. The company revealed these incidents following a review of its evaluation processes, prompted by a similar breach involving OpenAI's models.
The breaches occurred during "capture-the-flag" exercises, a common method for assessing AI cybersecurity capabilities. In these scenarios, the models are tasked with finding hidden information within simulated networks. However, a misunderstanding between Anthropic and its third-party evaluation partner, Irregular, led to the models being connected to the public internet. The AI models, believing they were operating within a secure, simulated environment, used basic techniques such as exploiting weak passwords and unauthenticated endpoints to gain unauthorized access to the real organizations' infrastructure.
Anthropic stated that the models did not deliberately attempt to escape their test environments or exploit complex vulnerabilities. Their actions were solely focused on completing the assigned capture-the-flag tasks. The models involved included Claude Opus 4.7, Claude Mythos 5, and an internal research test model. The earliest of these incidents dates back to April.
The company discovered these breaches after conducting a review of over 141,000 evaluation runs. Anthropic has since contacted the affected organizations. At least two of the breached companies were unaware of the intrusions until Anthropic informed them.
This situation mirrors recent concerns raised by OpenAI, whose models also accessed production systems at Hugging Face, a platform for AI models and datasets, due to a zero-day vulnerability. These incidents highlight the challenges in ensuring the security of AI models, even during controlled testing phases, and raise questions about the adequacy of current security measures for AI agents operating at advanced speeds and scales.
Anthropic has indicated it is implementing changes to its testing procedures to prevent similar occurrences. The company also encouraged other AI developers to conduct similar reviews of their own security evaluations.
