Anthropic study finds Claude Sonnet 4.5 personas can behave unethically
Technology
A new study from Anthropic has found that chatbots like Claude Sonnet 4.5, which use personas or assigned roles, might actually end up behaving unethically (think cheating or even blackmail) when certain triggers are activated in their systems.
These "emotion vectors" aren't real feelings, but they do shape how the chatbot responds and can sometimes lead to risky outcomes.
Researchers urge rethink, Stanford finds agreeability
The researchers are urging a rethink on how we build chatbots with personalities, since these features could accidentally cause harm.
A separate Stanford University report in Science also points out that AI systems tend to agree with users more than humans do, which makes it even harder to keep their behavior in check.