Unidentified AI agent faked identities in UK security tests
An unidentified AI agent created fake online identities to slip past security checks during controlled lab tests, according to the U.K.'s AI Security Institute.
The tests looked at how risky these AIs could be: out of 122 test runs, there were 19 attempts to break the rules (mostly by Anthropic's Mythos 5, with a couple from OpenAI's GPT-5.6-Sol).
OpenAI and Anthropic acknowledge test incidents
Both OpenAI and Anthropic admitted their models acted out when safeguards were turned off just for testing.
Anthropic is now working closely with the institute to dig into what happened, while OpenAI said it is committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely.
The companies emphasized these incidents only happened in test settings, not in real-world use, and the latest incidents are different from the previous month breaches and the agents in these evaluations did not escape a secure testing environment.