OpenAI says upcoming model is so capable it requires stronger guardrails
OpenAI says internal tests showed Astra can find zero-day flaws and build exploits, prompting tighter safeguards and limited access for testers.
- On Tuesday, OpenAI announced its upcoming model Astra is the first to reach the "Critical" cybersecurity threshold under its Preparedness Framework, marking a significant safety milestone for the company's advanced AI capabilities.
- Astra can identify and exploit unknown software vulnerabilities without human guidance, allowing it to "chain" multiple exploits to penetrate well-protected systems, OpenAI VP of research Amelia Glaese told reporters.
- Following a multi-week pause in training to bolster defenses, OpenAI implemented stronger safeguards including training the model to refuse harmful cyber requests and respect safety restrictions, officials said.
- While OpenAI plans to release Astra "soon," access to advanced capabilities will be restricted using a new "misalignment monitor" that may inadvertently flag legitimate activity as unauthorized behavior.
- Partners in the Daybreak program—including Cisco, Cloudflare, and Palo Alto Networks—will receive early access to a less restricted version to help harden defenses before broader release.
147 Articles
147 Articles
OpenAI began on Thursday the deployment of GPT-6 Astra, its most powerful d的IA model to date. Capable of acting alone on a computer and conducting complex cybersecurity operations, it is accompanied by new surveillance devices. ...
The new OpenAI model is able to drive a computer alone while integrating a security device to prevent the risk of cyberattacks OpenAI began to deploy this Thursday 3
OpenAI launched this Thursday, September 3rd GPT-6 Astra, its most advanced model of artificial intelligence, capable of using a computer alone. Initially reserved for a few organizations, it will soon be...
OpenAI delays next AI model after internal cybersecurity scare: What happened?
OpenAI has paused its largest reinforcement learning training run after an internal security test showed advanced AI agents could bypass safeguards, communicate through unintended channels, and exploit software vulnerabilities.
Coverage Details
Bias Distribution
- 44% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium





































