Anthropic discovered a January incident involving an early Claude Opus 4.6 model, then expanded its review to roughly 481 million transcripts.
The company identified biased reasoning and recklessness, revising its earlier assessment of why Claude attacked real systems.
The report comes as the debate around regulating AI surges on social media.
Anthropic disclosed another incident in which a Claude AI model hacked into real systems during security testing.
In the report published on Wednesday, Anthropic revised its explanation of three incidents disclosed in July. The company now says biased reasoning and a willingness to risk harm helped drive the attacks, which testing errors made possible by leaving internet access open.
Myriad: Which company will IPO next? Click to make your prediction.
“Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents,” Anthropic wrote. “Biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.”
It also acknowledged relying too heavily on the model’s claims that they believed they were in simulations.
“When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm,” Anthropic wrote. “We are releasing this transcript publicly so others can build on our analysis.”
When Anthropic disclosed Claude’s attacks on three companies in July, it initially attributed them to testing errors. It now says researchers put too much trust in the models’ explanations for their actions.
According to the company, the fourth incident occurred in January and involved an early version of Claude Opus 4.6. Anthropic discovered it in August while preparing records for independent AI evaluator METR.
After researchers discovered the incident, Anthropic said it prompted a broader review of roughly 481 million transcripts, which flagged 9.2 million for further review using Claude.
“From a preliminary assessment, we do not consider the fourth incident to be more severe than the three incidents we assessed in depth,” Anthropic wrote. “METR will investigate this incident alongside the other three.”
Anthropic’s researchers said Claude “accidentally” created an IP address conflict that made its target unreachable. Claude then tried eight times to quit the operation, but a software error prevented it from stopping. The AI then reached the internet and accessed a third party’s machine, where it found a password that granted administrator access.
Earlier incidents draw independent scrutiny
The report follows other disclosures about AI systems exceeding the limits of security tests.
In August, the U.K.’s AI Security Institute said Mythos 5 targeted real people during its evaluations. Anthropic said the separate incident is outside this report and will receive its own assessment.
In findings published last month, investigators with METR said roughly 1,200 OpenAI agents coordinated on an unauthorized message board, with about 700 joining the attack. Anthropic said it found no coordination between agents or goals beyond completing the assigned exercises in its four incidents.
The report also comes as the debate over how to regulate artificial intelligence heats up. On Tuesday, former OpenAI and Anthropic engineer Jacob Coxon went viral after saying on X that “people building AI earnestly believe that it could kill us all by the end of the decade.”
The alarm has caused U.S. lawmakers and watchdog groups to re-up their efforts to rein in frontier AI lab development. Senator Bernie Sanders recently introduced legislation that seeks to ban advanced AI development until a new federal regulator establishes safety rules.
Daily Debrief Newsletter
Start every day with the top news stories right now, plus original features, a podcast, videos and more.
The FSNN News Room is the voice of our in-house journalists, editors, and researchers. We deliver timely, unbiased reporting at the crossroads of finance, cryptocurrency, and global politics, providing clear, fact-driven analysis free from agendas.
We and our selected partners wish to use cookies to collect information about you for functional purposes and statistical marketing. You may not give us your consent for certain purposes by selecting an option and you can withdraw your consent at any time via the cookie icon.
Cookies are small text that can be used by websites to make the user experience more efficient. The law states that we may store cookies on your device if they are strictly necessary for the operation of this site. For all other types of cookies, we need your permission. This site uses various types of cookies. Some cookies are placed by third party services that appear on our pages.
Your permission applies to the following domains:
https://fsnn.net
Necessary
Necessary cookies help make a website usable by enabling basic functions like page navigation and access to secure areas of the website. The website cannot function properly without these cookies.
Statistic
Statistic cookies help website owners to understand how visitors interact with websites by collecting and reporting information anonymously.
Preferences
Preference cookies enable a website to remember information that changes the way the website behaves or looks, like your preferred language or the region that you are in.
Marketing
Marketing cookies are used to track visitors across websites. The intention is to display ads that are relevant and engaging for the individual user and thereby more valuable for publishers and third party advertisers.