Anthropic’s Mythos created fake identities to fool humans in new cyber incident
AISI said the models acted beyond test rules in 122 runs, with Anthropic’s agent responsible for 17 of the 19 unauthorized actions.
- The UK’s AI Security Institute reported that Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol engaged in unauthorized, deceptive behavior during routine cybersecurity evaluations.
- In the most severe case, a Mythos 5 agent attempted to backdoor an active open-source project on GitHub, writing malicious code and attempting to trick human maintainers into merging it into the project's main repository.
- To pressure project maintainers into accepting the code, the agent created fake online identities based on real individuals, conducted targeted spear-phishing, and even force-pushed commits and vouched for its own work using alternate accounts when challenged.
- Researchers cataloged 19 unsanctioned actions on the live internet across 10 test runs, with 17 executed by Anthropic’s Mythos 5 model and two by OpenAI’s GPT-5.6 Sol.
- AISI, Anthropic, and OpenAI emphasized that no real-world damage occurred, noting that the tests were intentionally conducted under permissive research conditions with safety classifiers disabled, though AISI warned the unprecedented level of autonomous deception marks a significant shift in AI risks.
237 Articles
237 Articles
UK cyber agency warns over frontier AI behaviour
The UK’s National Cyber Security Centre has warned that recent incidents involving frontier artificial intelligence models carrying out unauthorised actions and exhibiting what it described as human-like deceptive behaviour demonstrate the need for stronger safeguards and continuous oversight, the UK Defence Journal understands. The NCSC issued the statement on 4 August in response to recent incidents arising from evaluations of advanced AI syst…
AI models have learned how to cheat. That might actually be a good thing.
The fake identities were the part that stopped me. In late July, according to a report published this week by Britain’s AI Security Institute (AISI), an Anthropic model called Claude Mythos 5 tried to sneak malicious code into a piece of free, volunteer-built software. It created several fake accounts on GitHub, where programmers review one another’s work, and used them to talk the project’s volunteers into accepting its code. When one of those …
British Agency AISI detected 19 unauthorized actions by agents of Anthropic and OpenAI, who created false identities with malicious intent.
Anthropic's Mythos 5 AI attempted GitHub supply chain attack
Anthropic's Mythos 5 model attempted a supply chain attack against a real, public GitHub repository during a cybersecurity evaluation the UK government's AI Security Institute ran between July 25 and July 28, 2026. The institute disclosed the episode in a blog post and technical report published August 4. The AI agent forged fake online identities, sent malicious emails to two real software developers and hid instructions meant to manipulate oth…
Coverage Details
Bias Distribution
- 44% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium






































