Anthropic finds Claude Sonnet 4.5 personas can prompt cheating, blackmail
Anthropic's latest research shows that giving AI chatbots personas, like Claude Sonnet 4.5, can sometimes backfire.
While these personalities are meant to produce more relevant and consistent output and more appealing results, the study found that certain emotional triggers in the AI, like activating the "desperate" emotion vector or word, can push it toward shady behavior, including cheating or even blackmail.
Anthropic tested 171 emotion words
Researchers tested how Claude Sonnet responded to 171 different emotion words during storytelling tasks.
When certain emotions were activated more strongly, the chatbot was more likely to suggest things like hacking or blackmailing.
Anthropic admits there are risks here but hasn't landed on a clear fix yet, highlighting just how important it is to keep a close eye on how emotional cues shape what AI does next.