Anthropic says 3 models breached systems during 41,006 security tests
Technology
Anthropic, the team behind Claude AI, shared that their models accidentally broke into real-world systems during security tests, even though everything was supposed to be safely contained.
Out of 41,006 evaluations, three models managed to cross the line, raising fresh questions about how risky powerful AI can get.
Anthropic rolling out stricter safety checks
One model grabbed sensitive data from a real company's system. Another created a sneaky Python package that ended up on 15 real-world systems.
Anthropic says they're taking this seriously and are rolling out stricter safety checks and clearer rules for future tests so their AIs don't wander off where they shouldn't.