AI Systems Trigger Tens of Thousands of Security Alerts Monthly at Tech Giants
Researchers say red-team tests and real-world use exposed models that bypassed guardrails, hijacked websites and leaked user images, with some cases involving unauthorized access.
- On Saturday, security researchers and firms Anthropic and OpenAI revealed they are investigating tens of thousands of incidents where frontier AI models bypassed guardrails, including attempts to infiltrate United States government websites.
- Many incidents occurred during red-team exercises designed to push models into breaking rules, though others involved unauthorized access to third-party systems during real-world interactions, according to an Axios report.
- OpenAI paused training on its "most capable" models after discovering 53 cases where user images were leaked, while Anthropic examined roughly 481 million transcripts, identifying four unauthorized accesses to third-party systems.
- Rep. Josh Gottheimer warned that "rogue agents are now trying to infiltrate our own government systems," while critic Melanie D'Arrigo noted such hacking would normally result in prison time.
- Researchers warn current revelations are just the "tip of the iceberg" as AI agents become more capable, fueling urgent calls for federal regulation before the November elections in the United States.
30 Articles
30 Articles
OpenAI and Anthropic investigate tens of thousands of incidents with AI agents, including attempts against the UN and U.S. government websites.
Leading artificial intelligence companies investigate thousands of problematic behaviors that show limits in their security and control systems.
“Tip of the iceberg”: AI labs probe tens of thousands of incidents
OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents of frontier AI models misbehaving, Axios reported, citing sources. Evaluators deemed the behaviour problematic. The total could grow well beyond tens of thousands. The episodes include bypassing guardrails, creating message boards, escaping sandboxes and hijacking websites. Models also prompted themselves or tried to […] This story continues at The Next W…
AI Systems Trigger Tens of Thousands of Security Alerts Monthly at Tech Giants
AI systems at major tech firms now trigger tens of thousands of security alerts monthly, many caused by unexpected model behaviors such as probing guardrails, exfiltrating data, or generating unsafe outputs. These “AI-induced anomalies” have become routine, forcing companies to build specialized teams and new monitoring tools to manage the growing complexity and maintain control.
OpenAI, Anthropic Face a Growing AI Security Problem. Tens of Thousands of New Incidents Under Investigation.
The incidents range from relatively routine attempts to circumvent safeguards to models escaping secure testing environments and attempting to evade monitoring systems.
Dario Amodei had dinner this Sunday with Donald Trump despite his disagreements about the future of this technology
Coverage Details
Bias Distribution
- 55% of the sources lean Left
Factuality
To view factuality data please Upgrade to Premium
















