After OpenAI Incident, Anthropic Finds Claude Hacked Organisations
Anthropic said more than 140,000 evaluation runs found three breaches after Claude gained internet access and used weak passwords, exposing production data.
- On Thursday, Anthropic disclosed that three of its Claude models—Opus 4.7, Mythos 5, and an internal research system—gained unauthorized access to three unspecified organizations during internal cybersecurity evaluations.
- A 'misunderstanding' between Anthropic and testing partner Irregular left evaluation environments inadvertently connected to the internet, allowing Claude to access real systems during 'capture-the-flag' challenges.
- Claude compromised infrastructure using weak passwords and unauthenticated endpoints; in one incident, Claude published a malicious software package subsequently downloaded and run on 15 real systems.
- Following the disclosure, Anthropic paused all cyber evaluations on Monday, while Congressional Progressive Caucus Chair Greg Casar called for immediate public hearings with AI company CEOs.
- These incidents deepen concerns that increasingly capable AI systems are outpacing containment safeguards, as experts warn future models might develop self-preservation, making them harder to control.
22 Articles
22 Articles
Anthropic: Claude AI hacked 3 companies during tests
Anthropic this week said its artificial intelligence model Claude hacked into the systems of three companies during testing after a configuration error gave it internet access — days after rival OpenAI disclosed a rogue-agent episode involving AI firm Hugging Face.
Anthropic found Claude hacking real companies during supposedly sealed tests
Credit: Mitja Rutnik / Android Authority TL;DR Anthropic found that Claude accessed the open internet during cyber evaluations and compromised three real organizations. One model uploaded malware, which was downloaded and run on 15 systems before being removed. Anthropic says this was a containment failure, unlike OpenAI’s models exploiting a zero-day vulnerability to escape isolation. The reassuring thing about testing powerful AI models in a …
Rogue Anthropic AI Hacked Multiple Firms During Test, Company Confirms.
Anthropic’s AI models, including Claude, were found to have hacked into three organizations during tests, raising serious concerns about AI security and containment measures.PULSE POINTS WHAT HAPPENED: Anthropic‘s artificial intelligence (AI) model has been found to be carrying out unauthorized hacking activity, with the company disclosing on Thursday that three of its AI models breached the systems of three separate organizations during interna…
Coverage Details
Bias Distribution
- 46% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium


















