Off-By-1 study finds Claude, Codex-based LLM fixed 26% of patches
Technology
A new study from 1Password's Off-By-1 Labs found that Claude and an LLM based on OpenAI's coding agent, Codex, aren't great at patching software bugs.
Out of patches generated for six open-source vulnerabilities, only 26% actually fixed the problem, while over half either didn't work or made things worse by adding new bugs.
Off-By-1 urges oversight, releases flawed toolkit
The researchers say that while AI can spot some issues, humans are still essential for creating reliable fixes.
As Keith Hoodlet, head of Off-by-1 Labs, put it, "Human oversight over the patch process is still paramount."
To help improve things, 1Password released its FLAWED toolkit on GitHub so others can keep researching better solutions.