You are connecting from Lake Geneva Public Library, please login or register to take advantage of your institution's Ground News Plan.
Published 7 hours ago • loading... • Updated 1 hour ago
OpenAI says upcoming model is so capable it requires stronger guardrails
On Tuesday, OpenAI announced its upcoming model Astra is the first to reach the "Critical" cybersecurity threshold under its Preparedness Framework, marking a significant safety milestone for the company's advanced AI capabilities.
Astra can identify and exploit unknown software vulnerabilities without human guidance, allowing it to "chain" multiple exploits to penetrate well-protected systems, OpenAI VP of research Amelia Glaese told reporters.
Following a multi-week pause in training to bolster defenses, OpenAI implemented stronger safeguards including training the model to refuse harmful cyber requests and respect safety restrictions, officials said.
While OpenAI plans to release Astra "soon," access to advanced capabilities will be restricted using a new "misalignment monitor" that may inadvertently flag legitimate activity as unauthorized behavior.
Partners in the Daybreak program—including Cisco, Cloudflare, and Palo Alto Networks—will receive early access to a less restricted version to help harden defenses before broader release.
"Astra is able to detect more security gaps than previous models. It is the first time that a model triggers the most stringent, so far only theoretical security requirements of OpenAI.
OpenAI is preparing an Astra, the first model of a company that has reached the critical threshold of cybersecurity and has been able to detect two self-identifiers of the "zero-day" type.