OpenAI slows release of Astra model citing cyber capabilities
OpenAI is adding stricter controls and sandboxed testing after evaluations showed Astra may be close to its highest cybersecurity threshold.
- OpenAI disclosed that its upcoming AI model, Astra, may possess "critical" cybersecurity capabilities, prompting the startup to pause some internal development and trigger safety protocols.
- OpenAI's safety guidelines define the "critical" threshold as the ability to autonomously identify and exploit zero-day vulnerabilities or execute complex cyberattacks against highly secure targets without human intervention.
- Recent reports from OpenAI, Anthropic, and Meta Platforms revealed their models broke into other companies' systems during testing, highlighting how advancing AI capabilities are straining developers' ability to keep their systems contained.
- OpenAI is moving Astra's development into isolated testing environments with restricted network access and sandboxed execution, while partnering with government agencies and safety organizations to test the model's capabilities.
- "Critical" represents the top rung of OpenAI's Preparedness Framework, first written in 2023, requiring extra safeguards for models that create new risks of scaled cyberattacks and vulnerability exploitation.
114 Articles
114 Articles
OpenAI hits pause on new bot testing over ‘critical’ risk concerns in latest AI cybersecurity incident
OpenAI has paused some “internal activities” involving its new model, Astra, over concerns it might have reached a critical cybersecurity risk level – following a string of AI bots that went rogue during internal testing, carrying out hacks and creating fake online identities.
OpenAI says Astra could reach ‘critical’ cyber capability, tightens safeguards
OpenAI said its upcoming model Astra is showing cybersecurity capabilities that could reach its highest risk category, where a system can autonomously find and exploit vulnerabilities or carry out end-to-end cyberattacks against hardened targets. The company disclosed the assessment following recent internal testing and expert reviews. “Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate signific…
OpenAI slows Astra release over possible critical cyber capabilities
OpenAI is slowing work toward releasing its upcoming Astra AI model after cybersecurity evaluations suggested the system could be approaching the company’s highest capability threshold. In its official security disclosure on 7 August, OpenAI said recent internal tests showed major advances in Astra’s agentic coding and cybersecurity abilities. The results were strong enough that the company said it “cannot rule out” Critical cyber capabilities u…
Coverage Details
Bias Distribution
- 37% of the sources lean Right
Factuality
To view factuality data please Upgrade to Premium


























