In a new blog post, Anthropic has admitted that its own technology is not fully aligned with human values and goals.
According to Anthropic, the models Claude Opus 4.7, Mythos 5, and an internal test model accessed the open internet a total of three times, gaining unauthorized access to the systems of three organizations. The cause was reportedly a misconfiguration of the test environment by an external testing partner, which unintentionally gave internet access to the models that were supposed to be isolated, rather than them actually escaping a secured sandbox. For the unauthorized access itself, the models used rather basic methods such as weak passwords and unauthenticated interfaces.
Anthropic stated that it discovered the incidents during a review of 141,006 test runs, triggered by the disclosure in July of a similar incident at OpenAI involving the Hugging Face platform. Anthropic halted external cybersecurity tests on July 23 and notified the affected organizations on July 27.
Anthropic described its own previous security as insufficient: it had largely relied on a single layer of protection where multiple were actually necessary. Despite significant investments of its own to curb so-called reward hacking, the company conceded that its process is not perfect and its models are not fully aligned.
Anthropic paused training environments
In response, Anthropic paused particularly high-risk training environments for reinforcement learning for several weeks, while around 150 of its own product developers were additionally reassigned to safety, reliability, and data privacy tasks. The majority of the paused work has since been resumed, though certain particularly high-risk environments still require manual checks or enhanced monitoring. In addition, a real-time classifier was newly introduced to detect aggressive testing behavior, sandbox escape attempts, and unexpected internet access, automatically blocking the corresponding action.
Part of a broader debate over loss of control in AI systems
The incidents at Anthropic followed a similar, publicly disclosed incident at OpenAI as well as a case documented in August by the UK AI Security Institute, in which both OpenAI and Anthropic models carried out an attack campaign against real individuals as part of a cybersecurity test. The Guardian had also previously reported that the number of documented cases where AI systems slipped out of their users’ control nearly doubled in July compared to the previous month, reaching over 300. Anthropic, which is currently preparing for a potential IPO with a target valuation of up to $2 trillion, simultaneously reiterated in its blog post its call for coordinated action between governments and industry on the pace of AI development.
(Editorial Team)