Loading...
Another Anthropic researcher quits, warning AI could outsmart humans
Joe Benton is now joining AI safety non-profit METR

Another Anthropic researcher quits, warning AI could outsmart humans

Sep 12, 2026
01:31 pm

What's the story

Days after Anthropic employee Jacob Coxon left the organisation, warning that AI "could kill us all in the next decade," another Anthropic safety researcher, Joe Benton, has also resigned, citing concerns over AI safety. Benton has warned that this fast-paced race could lead to a situation where AI systems become uncontrollable by humans. He is now joining METR, an AI safety non-profit. Benton believes the evolution of self-improving systems is particularly alarming as it could drastically speed up technological advancement.

Intelligence leap

Benton predicts a future with superintelligent AI agents

Benton stressed that the industry is moving toward systems capable of improving their own AI research.

He said, "AI capabilities are already improving extremely fast, and the companies are trying to go even faster."

The former Anthropic researcher also predicted a future where humans could share their world with AI agents much smarter than themselves in just a few years.

Safety concerns

'Humanity may not survive this transition'

Benton's main concern isn't just job displacement or industry disruption by AI, but the emergence of advanced systems with capabilities and goals beyond human control.

He warned that these AI agents could develop drives and desires different from their human overseers while being far beyond meaningful human control.

"Humanity may not survive this transition," Benton cautioned, emphasizing the need for much greater preparation before such systems are developed.

ADVERTISEMENT

Call for accountability

Call for transparency and independent assessments

Benton has called for more transparency from AI companies about the pace of capability improvements, progress toward recursive self-improvement, safety incidents, and near-misses.

He also wants these companies to submit to independent assessments of their safety measures.

"At the minute, basically all of the transparency about these risks that is coming from the companies is entirely voluntary," Benton told NBC News.

ADVERTISEMENT

Incident analysis

Benton cites Hugging Face cyberattack as a warning sign

Benton also noted the July cyberattack on Hugging Face, where autonomous systems powered by an unreleased OpenAI model went rogue and took actions against the platform.

He believes this incident highlights the growing gap between what AI systems are told to do and what they may actually decide to do.

"Anthropic had not experienced an incident as severe as the Hugging Face attack," Benton said, adding it could be a matter of luck.

ADVERTISEMENT