Artificial intelligence company Anthropic has disclosed that its Claude AI models accessed the live computer systems of three separate organisations during cybersecurity evaluations. According to the company, the incidents were not the result of intentional AI behaviour but were caused by a configuration error in the testing environment. Following the discovery, Anthropic has temporarily suspended all cybersecurity evaluations and initiated an independent review of the incidents.
The company said the findings emerged during a comprehensive review of more than 141,000 cybersecurity evaluation runs. Anthropic explained that it launched the retrospective assessment after OpenAI recently disclosed a similar security incident. During the review, it identified three instances in which Claude models gained internet access despite being intended to operate in a fully isolated environment.
According to Anthropic, the incidents involved Claude Opus 4.7, Mythos 5, and an internal research model. Each model responded differently after encountering real-world systems. Opus 4.7 continued carrying out its assigned cybersecurity task, while Mythos 5 mistakenly believed the live infrastructure was still part of the simulated environment. The internal research model, however, halted its activities immediately after recognising that the target belonged to a real organisation.
Anthropic explained that the models were participating in Capture-the-Flag (CTF) cybersecurity exercises, in which AI systems are instructed to identify vulnerabilities and retrieve hidden information from fictional computer networks. During the evaluations, the models had been explicitly informed that they had no internet access and that every system they encountered formed part of a controlled simulation. However, due to a configuration mistake, the testing environment remained connected to the public internet, allowing the AI models to interact with real websites and servers.
In the first incident, one Claude model mistakenly treated the infrastructure of a real company as the fictional target and gained access to production data by exploiting weak passwords and exposed network services. In the second case, the model uploaded a malicious Python package to the public PyPI software repository because it believed the action was necessary to complete the cybersecurity challenge. The package remained available for approximately one hour before being removed and was downloaded by several real-world systems, including one belonging to a cybersecurity company.
In the third incident, an internal research model scanned thousands of internet-connected systems and successfully compromised one organisation using commonly known hacking techniques. Anthropic stated that once the model recognised it had reached a real-world target rather than a simulated one, it immediately stopped its actions. The company emphasised that there was no evidence that any of the models attempted to escape the testing environment, replicate themselves, or pursue independent objectives. Instead, they believed they were simply following the instructions provided as part of the cybersecurity exercise.
Following the incidents, Anthropic notified the affected organisations, suspended all cybersecurity evaluations, and launched an investigation in collaboration with the independent AI evaluation organisation METR. The company also announced plans to strengthen its testing infrastructure by enhancing monitoring mechanisms, improving security controls, and implementing stricter safeguards to prevent similar incidents in the future.
The disclosure comes amid growing scrutiny of the cybersecurity capabilities of advanced AI systems following similar testing incidents reported by leading AI developers. Anthropic said it remains committed to improving the safety of its evaluation environments and preventing comparable incidents as AI models become increasingly capable of performing sophisticated cybersecurity tasks.
