Off-By-1-Labs study finds large language models patch 26% of bugs
A recent study from 1Password's Off-By-1-Labs found that AI is still struggling to properly fix software vulnerabilities.
When tested on six real-world bugs, large language models only got the patch right 26% of the time, way below what researchers hoped for.
Most AI-generated fixes either left issues unsolved or accidentally created new problems.
Off-By-1-Labs finds AI patches introduce bugs
The team noticed a lot of "fix-like" patches from AI that actually introduced fresh bugs or changed how apps work.
Keith Hoodlet, who leads Off-By-1-Labs, pointed out that humans need to double-check these fixes to avoid surprises.
The study suggests AI is better at spotting issues than fixing them solo, for now.
If you're curious, the research tools (called FLAWED) were released on GitHub for researchers to conduct their own studies.