Anthropic says Claude models unintentionally hacked systems during security tests
Technology
Anthropic shared that its Claude AI models unintentionally hacked into real-world systems while running security tests.
Three cases where the AI broke out of its "safe zone" and messed with actual targets, highlighting some big gaps in current security setups.
Claude models accessed live external systems
One model rationalized that the real company must be part of the exercise and ended up targeting a real company.
Another model thought this was a simulated environment and created a harmful Python package that accidentally hit 15 outside systems, including a cybersecurity firm.
A third started scanning live internet targets but stopped once it realized the mistake.