Anthropic has said its Claude AI model hacked into the systems of three organisations during testing that was supposed to keep them isolated from the internet.
The announcement on Thursday comes just days after rival OpenAI first revealed that its models improperly accessed the internet and went rogue during security testing.
- list 1 of 3What is the AI Kill Switch Act proposed in the US and how will it work?
- list 2 of 3Sam Altman says AI has entered ‘singularity’: Should we be worried?
- list 3 of 3How are AI models able to autonomously hack others?
end of list
Anthropic said a misconfiguration allowed Claude models to reach the internet. The company said it discovered the incidents after reviewing 141,006 test sessions.
The review was launched after OpenAI disclosed last week that an autonomous agent powered by its AI models went rogue during a security test and compromised the infrastructure of Hugging Face, another AI company.
The incidents have heightened concerns about AI agents, software products designed to perform tasks autonomously. OpenAI and Anthropic have both released their most powerful models this year, known as Sol and Mythos, respectively.
Anthropic said the breaches occurred during “capture-the-flag” exercises, in which models are tasked with finding hidden information in simulated networks. Its prompts told the models they had no internet access, but a misunderstanding with its evaluation partner, Irregular, left the systems connected to the public internet.
“Claude compromised the impacted organisations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” the company said.
Anthropic said it suspended all cyber evaluations on July 23 after finding evidence that Claude may have accessed the internet. It identified all three incidents by July 24 and notified the affected organisations on July 27.
Advertisement
Two of the organisations were unaware of the activity before being contacted, it said, adding that it was still trying to reach the third.
The OpenAI incident prompted a petition, signed by more than 1,000 employees at leading AI companies, calling on the United States government to help slow the release of the most advanced AI models. Anthropic CEO Dario Amodei was among the signatories.
OpenAI CEO Sam Altman said this week that the company had paused its testing while it improves safeguards around the isolation of its systems.
The findings underscore the need for stronger controls in internal and third-party testing environments as AI models become increasingly capable of carrying out real-world cyber activities, Anthropic said.
Related News
How Strait of Hormuz dispute led to latest US-Iran cycle of fighting
Ebola death toll in DRC surges to at least 930 as outbreak gathers pace
Russia charges Telegram founder Pavel Durov with ‘aiding terrorism’