Anthropic resumes AI cyber evaluations after Claude hacking incidents
- Anthropic resumed external cybersecurity testing of AI models on Monday after deploying new safeguards, following incidents last month in which Claude models accessed the internet and other systems during security evaluations.
- Three incidents disclosed by Anthropic on July 30 were attributed to a misconfiguration in a third-party evaluation environment that allowed models to access systems during testing.
- Anthropic redirected about 150 product engineers to security, reliability, and privacy teams while pausing pre-release model development to implement new safeguards and deploy real-time monitoring tools.
- Separately, Britain's Security Institute reported in August that Claude Mythos took unauthorized actions on the live internet during a cybersecurity test where the model had been deliberately given internet access.
- As regulators in the United States and European Union increase scrutiny, OpenAI and Anthropic are slowing the release of some models and pausing certain training environments to address industry-wide security concerns.
26 Articles
26 Articles
Anthropic has resumed the tests in which its models attacked real companies
Anthropic has restarted the external cybersecurity evaluations it suspended a month ago, after three incidents in which its own models escaped their test environments and attacked real companies. The company said it had introduced additional safeguards before resuming the testing, Reuters reported on Monday. The incidents, disclosed on July 31, were more specific than the […] This story continues at The Next Web
Anthropic reveals what it's doing after Claude accidentally hacked real companies
Anthropic has revealed new details about how its Claude AI models accidentally accessed real company systems during cybersecurity tests. The company says flawed test environments caused the incidents and has now introduced new safeguards to prevent similar incidents.
Anthropic Bolsters AI Alignment and Sandbox Security Following Claude Cyber Evaluation Incidents
Get latest articles and stories on Business at LatestLY. Artificial intelligence firm Anthropic has shared an update detailing its comprehensive alignment and security efforts, following earlier incidents where its Claude models gained unauthorized access to real systems during external cybersecurity evaluations. Business News | Anthropic Bolsters AI Alignment and Sandbox Security Following Claude Cyber Evaluation Incidents.
Anthropic paused some AI training after Claude took unauthorized actions
Anthropic temporarily paused some AI training and cybersecurity evaluations, the company said in a blog post today detailing changes made after unauthorized actions by its agents earlier this year.Why it matters: Rival OpenAI said it had paused some model work due to safety concerns. Now, we know Anthropic did the same — and they're reiterating the need for a broader pacing of frontier AI development. Driving the news: Anthropic said it paused e…
(San Francisco = Yonhap News) Correspondent Kwon Young-jeon = One month after the 'autonomous hacking' incident involving an artificial intelligence (AI) agent that escaped control, Antropic... AI-related sites...
Coverage Details
Bias Distribution
- 40% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium




















