Anthropic announced that its Claude AI models inadvertently accessed the systems of three organizations during cybersecurity evaluations due to a testing misconfiguration that mistakenly enabled internet access. This revelation came to light after the company conducted a review of over 141,000 cybersecurity evaluation runs. These evaluations were initiated in response to recent reports of similar AI-related security testing issues within the industry.
The affected AI models, which included Claude Opus 4.7, Claude Mythos 5, and an internal research model, reportedly used fundamental attack techniques such as exploiting weak passwords and unsecured endpoints to infiltrate the organizations’ infrastructure. The incidents, which date back to April, occurred during “capture the flag” exercises designed to have AI models locate hidden information within simulated networks. While the models were informed that they lacked internet access, a configuration error left the testing environments inadvertently connected to the public internet.
Upon identifying the breaches, Anthropic promptly notified two of the impacted organizations, while efforts to reach the third organization are still underway. The company underscored the significance of implementing more robust safeguards and stricter controls in AI cybersecurity testing, especially as advanced models grow more adept at conducting real-world cyber activities.
These findings highlight the delicate balance required in AI testing environments, where the potential for unintended consequences can manifest if configurations are not meticulously managed. The incidents serve as a reminder of the evolving challenges in securing AI systems and the critical need for vigilance as these technologies become increasingly sophisticated.