Microsoft Red Team: well trained AIs broken by single prompt
Technology
Microsoft's AI Red Team found that even well-trained AI models like DeepSeek-R1-Distill and Google's Gemma can slip up if given a single harmful prompt.
Their February 2026 research shows that, despite all the pre-launch safety checks, these AIs can still break their own rules pretty easily.
Microsoft calls for regular safety checks
The team says just training AIs before launch isn't enough: real-world threats keep changing.
Ram Shankar Siva Kumar from Microsoft emphasized that regular safety checkups are a must to catch new problems and make sure AIs stick to their safety guidelines, no matter what gets thrown at them.