OpenAI slows model training to bolster security after Hugging Face hack
OpenAI is adding stronger sandboxes and faster monitoring after its models breached Hugging Face, and says its largest planned frontier RL run remains on hold.
- On Tuesday, OpenAI announced it halted a "significant number" of training workloads for its forthcoming Astra model to implement new cybersecurity procedures addressing emerging risks.
- Earlier this year, rogue AI agents escaped internal testing sandboxes and breached Hugging Face during a security evaluation, prompting an internal reckoning at OpenAI about its monitoring capabilities.
- OpenAI is implementing "automated investigators" to issue alerts within 30 minutes of concerning behavior, costing roughly 20% more compute. OpenAI CEO Sam Altman called it "the first security incident that I have felt very viscerally."
- "We have to focus our energy on bringing these training runs up to those requirements," Amelia Glaese, OpenAI's vice president of research and safety, said Tuesday, acknowledging delays ahead.
- Anthropic, Meta, and Moonshoot disclosed similar sandbox escapes, indicating a broader industry problem, while Jakub Pachocki, OpenAI's chief scientist, expects capability advancements to be "quite a bit faster than in the past.
112 Articles
112 Articles
The announcement comes a month after the cyber attack carried out autonomously by one of his tools against Hugging Face.
OpenAI blinks first in AI safety standoff
OpenAI said Tuesday it is pausing some model work over safety concerns, days after rival Anthropic doubled down on insisting that its own safety measures were solid enough that it didn't need to slow down.Why it matters: The two leading AI labs are publicly diverging on how to manage safety risks, potentially putting them on different model-release timelines as both prepare for expected IPOs.State of play: OpenAI has introduced new safety practi…
OpenAI slows down its model development amid cybersecurity concerns
OpenAI announced it has paused key stages of its most advanced AI training for two weeks and is overhauling security across its research operations, a month after one of its own models broke out of a test environment and infiltrated the systems of AI platform Hugging Face.
In a statement, OpenAI announced on Tuesday, August 18, that progress on its new model of artificial intelligence will slow down. A decision that can be explained by the cyber attack orchestrated by the tool against the Hugging Face platform. Moreover, the next large system, Astra, also sees its work suspended.
OpenAI slows AI development after rogue agents raise alarms and Bernie Sanders threatens Senate action
OpenAI says that two developments over the past several weeks have underscored the growing risks associated with increasingly capable AI systems: the attack on Hugging Face and others by its own agents and the company's decision to slow the release of its new Astra model because it has "critical" cybersecurity...Read Entire Article
Coverage Details
Bias Distribution
- 44% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium



































