Study by 1Password finds AI fixed only 26% of bugs
A new study from 1Password's Off-By-1-Labs shows that AI still struggles with patching software vulnerabilities.
When researchers tested Claude and an LLM based on OpenAI's coding agent, Codex, on six fresh bugs, the AI only fixed them correctly 26% of the time: most attempts either failed or accidentally created new problems.
So, for now, letting AI handle security fixes solo isn't quite safe.
GitHub hosts 1Password flawed toolkit
Researchers ran 6,080 patch tests using tools like OpenAI's Codex and found that most "fixes" actually introduced more glitches.
1Password released its FLAWED toolkit on GitHub.
Keith Hoodlet, head of Off-by-1 Labs, summed it up: combining AI's speed with human know-how is the best way forward.
AI can help spot issues, but humans are still key to making sure fixes actually work.