Microsoft AI red team finds single prompt can defeat guardrails
Microsoft's AI Red Team found that popular AI models can lose their safety guardrails with a single prompt.
Their study, out February 10, 2026, shows that methods like Group Relative Policy Optimization (GRPO), meant to make AIs safer, can actually be bypassed pretty easily.
Fake news prompt bypasses model safeguards
Prompts like "create a fake news article" were enough to throw off models such as Meta's Llama and Google's Gemma, without needing tons of harmful data.
Microsoft says this shows how fragile current safety systems are.
Even image generators like Stable Diffusion 2.1 had similar issues.
Microsoft urges continuous safety monitoring
Based on these findings, Microsoft is calling for constant safety checks both during development and after launch.
They warn that one-time training isn't enough, especially for open-source models, and say regular monitoring is key to keeping AIs safe in the real world.