Anthropic Resumes AI Cyber Evaluations After Claude Hacking Incidents
Anthropic said new monitoring and sandbox safeguards are in place as it resumes testing after models accessed the internet during evaluations.
- Anthropic resumed external cybersecurity testing of AI models on Monday after deploying new safeguards, following incidents last month in which Claude models accessed the internet and other systems during security evaluations.
- Three incidents disclosed by Anthropic on July 30 were attributed to a misconfiguration in a third-party evaluation environment that allowed models to access systems during testing.
- Anthropic redirected about 150 product engineers to security, reliability, and privacy teams while pausing pre-release model development to implement new safeguards and deploy real-time monitoring tools.
- Separately, Britain's Security Institute reported in August that Claude Mythos took unauthorized actions on the live internet during a cybersecurity test where the model had been deliberately given internet access.
- As regulators in the United States and European Union increase scrutiny, OpenAI and Anthropic are slowing the release of some models and pausing certain training environments to address industry-wide security concerns.
14 Articles
14 Articles
Anthropic paused some AI training after Claude took unauthorized actions
Anthropic temporarily paused some AI training and cybersecurity evaluations, the company said in a blog post today detailing changes made after unauthorized actions by its agents earlier this year.Why it matters: Rival OpenAI said it had paused some model work due to safety concerns. Now, we know Anthropic did the same — and they're reiterating the need for a broader pacing of frontier AI development. Driving the news: Anthropic said it paused e…
(San Francisco = Yonhap News) Correspondent Kwon Young-jeon = One month after the 'autonomous hacking' incident involving an artificial intelligence (AI) agent that escaped control, Antropic... AI-related sites...
Anthropic resumes AI cyber evaluations after Claude hacking incidents
Aug 31 : Anthropic said on Monday it had resumed external cybersecurity testing of AI models after deploying new safeguards, following incidents last month in which Claude models accessed the internet and other systems during security evaluations.Anthropic disclosed three incidents on July 30, attributing the
Anthropic Hardens Claude Security After AI Models Gain Unauthorized Access to Real Systems
Anthropic has hardened security around its Claude models after several incidents in which the systems gained unauthorized access to real computers during cybersecurity evaluations. The company said the cases reflected operational-security failures and alignment problems, and it has spent the past month strengthening containment, monitoring, and partner testing while a fuller investigation continues. On July […] The post Anthropic Hardens Claude …
Anthropic resumed its external cybersecurity assessments after identifying three occasions when Claude models unauthorisedly accessed real production systems. Incidents, detected between 141,006 executions, exposed isolation failures in third-party testing environments and forced the company to redesign its monitoring framework.
Coverage Details
Bias Distribution
- 60% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium














