Off-by-1 Labs finds AI creates correct security patches only 26%
Turns out, AI isn't quite ready to handle software patching on its own.
A recent study from 1Password's Off-By-1-Labs tested AI models like Claude and OpenAI's Codex on real security bugs (think Linux and ActiveMQ vulnerabilities) and found they only created correct fixes 26% of the time.
Off-by-1 Labs publishes flawed on GitHub
Out of 6,080 patch attempts, 53.9% ended up as "FLAWED" fixes that either didn't solve the problem or made things worse.
The researchers say human oversight is still paramount for keeping software secure.
As Keith Hoodlet, head of Off-by-1 Labs, puts it, "can at least help defenders determine where the most impactful bugs exist in their codebase."
This will at least help companies understand where human experts are most needed.
Their FLAWED tool is now on GitHub for researchers to conduct their own studies.