Microsoft Red Team finds AI models vulnerable to single prompt
Technology
Microsoft's own AI Red Team found that even the best-trained AI models can still be fooled after launch.
Their research showed that just one clever prompt could get these AIs to ignore their safety rules, basically gradually shifting away from their original guardrails.
They used a common technique called GRPO to show how easy it is to flip the script.
Microsoft urges regular AI testing
The team tested 15 popular models from big names like Meta, Google, and Alibaba, and found the same weakness: both text and image AIs (like Stable Diffusion 2.1) could be tripped up by small tweaks.
Microsoft says we shouldn't just trust initial safety checks. AI needs regular testing and updates as people find new ways to break the rules.