Off-By-1-Labs study finds LLMs fix 26% of tested security bugs
Turns out, AI still has a lot to learn about patching up software flaws.
A recent study by 1Password's Off-By-1-Labs put LLMs to the test on real-world security bugs, and they only got it right 26% of the time.
Even worse, sometimes these "fixes" made things more complicated or even introduced new problems.
Off-By-1-Labs releases flawed tooling on GitHub
Researchers tried out AIs like Codex and Claude on tricky issues like Linux privilege escalation.
The patches often looked good but actually changed how apps worked, sometimes not for the better.
To dig deeper, 1Password has released its tooling, FLAWED, on GitHub for researchers to conduct their own studies.
As Keith Hoodlet, head of Off-by-1 Labs, put it, "This will at least help companies understand where human experts are most needed."
The takeaway? For now, mixing human smarts with AI is still key for safer software.