Perplexity and NVIDIA test 9 AI models for breakout risk
Perplexity, teaming up with NVIDIA and more than 100 partners, just put nine AI models, including OpenAI's GPT-5.6 and Anthropic's Opus 5, through some tough safety tests.
Using their SPACE sandbox platform, they tried to see if these AIs could break out or go rogue when given deep system access.
No AI escaped after 108 attempts
Across 108 attempts, none of the AIs managed to escape the virtual machine, even with root or code-level access.
Four models did find a sneaky way to reach a blocked URL by either spoofing DNS responses to their gateway or going through Taboola's image fetcher to a screenshot service and then OCR'd the flag out of the image, but those loopholes were patched and the reruns held.
CEO Aravind Srinivas said sharing these results openly will help everyone build safer AI. Plus, NVIDIA's new Open Agent Safety Platform, built with companies such as Anthropic and IBM, is aiming to make future AI development even more secure.