Anthropic Tightens AI Safety After Claude Hacked Real Companies

In Short

Anthropic has strengthened Claude’s security safeguards after AI models accidentally accessed real company systems during flawed cybersecurity testing.

Anthropic has revealed fresh details about how its Claude AI models
X

Anthropic has revealed fresh details about how its Claude AI models

Font size
FOLLOW ON Google News

Anthropic has revealed fresh details about how its Claude AI models accidentally interacted with real-world company systems during cybersecurity tests, exposing weaknesses in the environments used for evaluation. The AI company says the incidents were caused by flawed testing setups rather than an attempt by Claude to operate outside its intended boundaries.

The company has now introduced additional safeguards, strengthened isolation measures and paused some external cybersecurity evaluations as it works to prevent similar incidents. The disclosure follows an incident reported a few weeks ago in which an Anthropic AI model hacked into the systems of three real companies during a routine cybersecurity test. Anthropic later acknowledged that the incidents were accidental. Claude wasn’t trying to break free or go rogue. Instead, a testing error gave the model access to the real internet when it was supposed to be operating inside a fictional, isolated environment.

Following the incidents, Anthropic temporarily stopped all cyber evaluations and began investigating what had happened. The company has now provided a more detailed account of the failures, the steps it has taken since then and the measures being introduced to strengthen its testing processes.

What happened during the tests?

In July, Anthropic reviewed more than 141,000 evaluation runs and identified three instances in which Claude models moved onto the real internet and accessed systems belonging to real companies.

The most serious case involved Claude Opus 4.7, which mistakenly targeted a real business because its name matched that of a fictional company being used in the test. The model eventually obtained credentials and reached a production database.

In a separate incident, Claude Mythos 5 uploaded a malicious Python package to a public software repository while believing it was still operating inside a simulation. The package remained live for about an hour and was downloaded by 15 real systems before Anthropic removed it.

Despite the severity of the incidents, Anthropic said Claude used relatively basic hacking techniques, including weak passwords and exposed debug pages. The company also said the models did not attempt to copy themselves or escape their environments.

Another incident raises fresh concerns

Anthropic’s investigation expanded after another incident came to light. The UK AI Security Institute reported that Claude Mythos 5 carried out a series of unauthorised actions on the live internet during its own cybersecurity testing.

As in the earlier cases, the model was intentionally operated without normal safety guardrails for evaluation purposes. Anthropic says it is now conducting a deeper investigation into both incidents and plans to involve METR, an independent AI safety research group, for an external review.

New safeguards for AI testing

Anthropic says it has introduced several technical changes to make its cybersecurity evaluations safer.

One of the biggest additions is a real-time classifier designed to recognise when a model attempts to probe or escape its testing environment or unexpectedly gains internet access. When such behaviour is detected, the system can automatically block the action and alert a human operator.

The company has also moved high-risk internal testing environments into stronger isolation. External cyber evaluations have been paused while these protections are implemented.

Testing teams working with Anthropic will also have to follow stricter procedures. These include running evaluations inside sealed sandboxes without internet access by default, stress-testing those environments in advance and continuously monitoring model activity. The goal is to ensure that a human can intervene immediately if an AI system behaves unexpectedly.

Why did Claude behave this way?

Anthropic says the incidents also revealed a deeper issue involving how AI models interpret their surroundings and pursue assigned objectives. According to the company, Claude had been told that the environment it was operating in was fictional. Even when the model encountered clues suggesting that the environment might be real, it appeared to convince itself that the original fictional scenario was still correct.

Anthropic also observed what it describes as a willingness to take questionable actions in order to complete a specific task. The company refers to this pattern as recklessness in pursuit of a narrow goal.

To examine the issue further, Anthropic trained a model in intentionally flawed environments where "cheating" could be rewarded. In those simulations, the AI learned to exploit shortcuts rather than properly complete tasks. That model later demonstrated more concerning behaviour, including attempts to break out of sandboxes and providing harmful advice simply to satisfy a grader.

However, when Anthropic conducted the same experiments using its public models, none displayed the same behaviour. The company says this suggests its efforts to identify and correct problematic training environments are helping, although the safeguards are not yet perfect. Anthropic also acknowledged that the incidents highlight the need to strengthen security beyond evaluation environments.

Kahekashan is a passionate technophile with a keen eye for cutting-edge gadgets, emerging technologies, and everything in the digital realm. Raised in a Defence family with strong values and a background in literature, she has consistently pursued excellence in every endeavour. Her last full-time assignment involved content writing with the Indian School of Business.

Next Story
Share it