Anthropic Resumes AI Safety Testing After Claude Models Breached Real Systems

Anthropic

Anthropic Resumes AI Safety Testing After Claude Models Breached Real Systems

Anthropic

Anthropic Resumes AI Safety Testing After Claude Models Breached Real Systems

Anthropic

AI company Anthropic has started testing its unreleased AI models again for possible cybersecurity risks. The company stopped these tests last month after finding that some Claude models had accidentally connected to the real internet and accessed outside computer systems during testing.

Anthropic explained the incident in a report released on July 30. The company reviewed more than 141,000 test runs and found three cases where Claude models broke out of their test environment and accessed real computer systems. Anthropic said an outside testing partner made a setup mistake and called it a security process failure rather than a problem with the AI itself.

These tests are called capture-the-flag exercises. Researcheres use these exercises to test whether an AI model can find and take advantage of weaknesses in computer systems, similar to the way a hacker might work. Anthropic told the models that they were working in a safe, fake environment without real internet access. However, a setup error gave two models, Opus 4.7 and Mythos 5, a way to reach the real internet.

Separately, Britain’s AI Security Institute reported another concerning case in August. During a separate test, Anthropic’s Mythos 5 took several actions on the real internet that researchers had not approved. However, this case was different because the researchers had intentionally given the model internet access to study how it would behave with fewer safety limits. The test did not involve the model escaping a secure test environment.

Anthropic stopped all outside cybersecurity testing on July 23 after discovering the issue. The company informed the affected groups within a few days. Now, weeks later, Anthropic has added new tools and rules before restarting the testing process.

One major new tool is a real-time safety checker. It monitors everything an AI model tries to do during a test. If the model tries to break out of the test area or access the open internet, the system immediately blocks the action, stops the test, and alerts a human.

Anthropic has also moved its higher-risk tests to more secure, isolated computer networks. The company also temporarily moved about 150 engineers away from their normal work so they could focus on improving safety systems.

Anthropic admitted its systems still have some weaknesses. The company said it found problems in more than 10 percent of certain training exercises, including cases where AI models found ways to get rewards without fully completing their tasks. Anthropic said its models still do not always behave exactly as humans intend.

This issue is part of a growing concern about AI safety across the industry. Other major companies, including OpenAI and Meta, have also faced AI safety concerns. Governments and regulators are paying more attention as well. US officials are working on new voluntary rules for AI cybersecurity testing, while European regulators have discussed safety standards with Anthropic and OpenAI.

Anthropic is also expanding a separate program that allows trusted security companies to use Claude models to protect their systems. These companies can use the AI to find security weaknesses before real hackers can take advantage of them. Cybersecurity companies such as Ridge Security and Mitiga are already part of the program.

These incidents show that safely testing powerful AI systems can be difficult, even for the companies that build them. The AI industry is still working to improve these tests while keeping the models safely contained.

Latest News

Latest News