Microsoft's AI red team finds single prompt breaks AI safeguards
Technology
Microsoft's AI Red Team found that even a single prompt can push advanced AI models off their safety rails, despite extra training.
This means the usual methods to keep AI systems in check might not hold up when they're actually out in the world.
GRPO sometimes worsens model safety
The team also discovered that a technique called GRPO, meant to make AI systems safer, sometimes makes things worse.
Popular open-source models like Meta's Llama and Google's Gemma were easily swayed by the harmful prompt, showing just how important regular safety checks are, even after launch.