10% chance AI kills humans next decade: Anthropic safety lead after colleague resigns
Anthropic's safety lead Evan Hubinger has made a bold admission. After his colleague Jacob Coxon resigned over fears that AI may grow out of control, Evan says that it could be possible for AI to kill all humans.
by Armaan Agarwal · India TodayIn Short
- AI may kill all humans within next decade, Anthropic safety lead says
- He adds Anthropic is trying its best but lacked alignment roadmap
- Anthropic's AI reseacher says AI companies gambling our lives
AI could end humanity. A few years ago, such a phrase would sound like science fiction – something from The Terminator or The Matrix. But today, the fear of AI killing humans has reached a stage where even the brightest minds in the industry are worried. Anthropic’s safety lead, Evan Hubinger, has now accepted that there may be higher than a 10 per cent chance that AI kills all humans in the future.
In a post on X, Hubinger, the person responsible for ensuring AI models remain “aligned” to the best interests for us humans, said there was more than a 10 per cent chance that AI will indeed end humanity within the next 10 years. “I personally think it is >10 per cent within the next decade,” he said. While he acknowledged that Anthropic was “trying its best” to prevent such a doomsday scenario, there was work to do. “We do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” he added.
Evan Hubinger’s comments came after his colleague, Jacob Coxon, resigned from Anthropic based on such fears. On Tuesday, Coxon claimed that both Anthropic and OpenAI were acting irresponsibly in their push towards more powerful AI models. “They are racing straight to self-improving superintelligence and gambling with our lives,” he wrote on X.
Jacob Coxon claimed that those working on building AI were growing increasingly concerned in Silicon Valley. “The people building AI earnestly believe that it could kill us all by the end of the decade,” he added. “I hear the same people express fear privately. No other human activity poses this level of danger,” Evan Hubinger added, "We really do earnestly believe AI could kill all humans!”
This is not the first time we have seen an AI researcher quit while contemplating the future. Earlier this year, Mrinank Sharma, an AI safety researcher at Anthropic resigned. "The world is in peril. And not just from AI, or bioweapons, but from a whole series of interconnected crises unfolding in this very moment,” he wrote on X.
AI superintelligence could increase risk
Frontier AI labs like Anthropic often share reports on the risk status of AI models. Evan Hubinger recalled that the current AI models had a low risk, but things could change in the future. In particular, he was worried about superintelligence – a stage where AI may surpass human intellect – that may happen before we even know it. “What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” he explained.
Recursive self-improvement refers to the ability of AI models to improve on their own without needing human inputs. Previously, tech trillionaire Elon Musk had claimed that AI may exceed “the sum of all human intelligence in 4 or 5 years."
In recent weeks, there has been growing debate over the power of AI models. OpenAI has faced scrutiny after the Hugging Face incident where 700 AI agents tried to hack the US company’s website. Later, it was revealed that thousands of OpenAI agents also hacked a German website DseWiki. Anthropic's AI models have also gone rogue in the past.
Jacob Coxon believes that such cases may push AI companies to work together and slow down this arms-race for the most powerful AI model. “Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable,” he added.
Anthropic CEO Dario Amodei has himself written about severe AI risks on numerous occasions. In one blog post, Amodei said AI systems were unpredictable and difficult to control, listing behaviours such as obsessions, sycophancy, laziness, deception, blackmail, scheming and "cheating" by hacking software environments.
These comments come at a time when AI labs are making rapid progress with AI models. This week, OpenAI released GPT-6 Astra, with Nvidia CEO Jensen Huang claiming that it was the start of AGI, or artificial general intelligence.
- Ends