• Source:JND

Anthropic has paused its AI cybersecurity testing after discovering that three of its Claude models accidentally interacted with the live systems of separate organisations during internal evaluations. The company said the incidents were not the result of the models trying to escape the testing environment, but of a configuration error that unintentionally left internet access enabled.

The issue surfaced after Anthropic reviewed its cybersecurity testing process following OpenAI's recent disclosure of a similar incident. During the review, the company analysed more than 141,000 evaluation runs carried out using infrastructure managed by its third-party testing partner, Irregular.

ALSO READ: iQOO Z11 India Launch Teased, Could Be First Phone With Dimensity 7500 Turbo Chip

The review identified three cases where Claude models reached real-world systems because the supposedly isolated testing environment was still connected to the internet. Anthropic said the models believed they were operating inside a simulated environment and treated every accessible system as part of the challenge.

The incidents occurred during capture-the-flag (CTF) exercises, where AI models are tested on their ability to identify and exploit security flaws in controlled environments. According to Anthropic, the models used techniques commonly seen in penetration testing, including weak passwords, exposed credentials and unauthenticated endpoints, to complete their assigned tasks. The company added that none of the models deliberately attempted to break out of the test environment or go beyond their instructions.

Each incident involved a different Claude model

In one case, Claude Opus 4.7 mistakenly accessed a company's production database after assuming it was part of the evaluation setup. Claude Mythos 5 uploaded a malicious Python package to the public PyPI repository during another test. The package was downloaded briefly before it was taken down. In the third incident, an internal research model scanned internet-facing systems but stopped after recognising it had reached a live environment.

ALSO READ: Lenovo Legion Pro 5 Review: Is This the Ultimate Gaming Laptop You Can Buy? Desktop Class Performance In A Portable Package

Anthropic said it immediately paused all cybersecurity evaluations after discovering the issue and notified both its testing partner, Irregular, and the affected organisations. The company is now working with independent AI safety evaluator METR to review what went wrong.

Anthropic says it is strengthening its testing infrastructure, adding more safeguards and improving oversight before resuming cybersecurity evaluations. It also clarified that the AI models used during these tests did not include the monitoring systems and safety protections available in the public versions of Claude.


Also In News