Study from Off-By-1-Labs finds AI 26% successful at fixing vulnerabilities
A new study from 1Password's Off-By-1-Labs found that AI tools like Claude and Codex still struggle to patch software vulnerabilities.
Out of thousands of tries on six real security flaws, AI only got it right about 26% of the time, and sometimes even made things worse by breaking features or adding new bugs.
Off-By-1-Labs: 6,080 patch attempts need oversight
The researchers ran 6,080 patch attempts and say that while AI can help spot problems, people are still needed to make smart decisions about fixes.
As Keith Hoodlet from Off-By-1-Labs put it,
Keith Hoodlet, head of Off-by-1 Labs, said, "This will at least help companies understand where human experts are most needed."
To help others dig deeper into how AI handles patches (and where it falls short), they've released a tool called FLAWED on GitHub for researchers to conduct their own studies.