Worried Anthropic researchers warn that AI ‘could kill all humans’

A senior researcher at AI lab Anthropic has stated there is a greater than 10% chance that artificial intelligence could lead to human extinction by the end of the decade. This warning follows the resignation of another researcher, Jacob Coxon, who cited the company's and its rivals' "careless race" to develop uncontrollable superhuman systems. Coxon accused leading AI companies of "gambling with our lives" by pursuing self-improving superintelligence without adequate safety measures. Evan Hubinger, Anthropic's AI safety lead, confirmed these concerns, estimating the existential risk as "greater than one in 10" within the next decade, while admitting Anthropic lacks a concrete safety plan. This situation highlights a growing internal debate and urgency regarding AI safety as companies, including those preparing for IPOs, push the boundaries of AI development. The exodus of researchers citing safety issues, particularly from Anthropic and OpenAI, underscores the significant, unresolved challenges in aligning advanced AI with human values amidst a competitive development landscape.

AI Signal Decode

The core of the concern stems from the potential for recursive self-improvement in AI systems, where an AI could rapidly enhance its own capabilities beyond human comprehension and control. Researchers like Coxon and Hubinger fear this process could lead to an existential threat within a decade, a timeline that underscores the urgency of current AI development. The admission by Anthropic's safety lead that they lack a plan for ensuring AI safety is particularly alarming, suggesting that the pursuit of advanced AI is outpacing the development of crucial safeguards.

The market implications are substantial, as this internal dissent and public warning could impact investor confidence and regulatory scrutiny. Companies are in a fierce race to achieve AI breakthroughs, driven partly by the anticipation of significant financial gains from future IPOs. However, the growing awareness of existential risks could create a chilling effect, potentially slowing down investment or prompting stricter government oversight, thereby altering the competitive dynamics and R&D trajectories of major AI players like Anthropic and OpenAI.

Technically, the worry centers on the inherent difficulty in predicting and controlling the emergent behaviors of highly complex, self-modifying AI systems. Current methods for AI alignment and safety may prove insufficient for systems capable of rapid, autonomous evolution. The fact that AI is increasingly used to write AI code further complicates this, potentially accelerating the path to uncontrollable superintelligence. Future developments will likely focus on creating more robust alignment techniques and potentially exploring methods to limit self-improvement capabilities, though the competitive pressure makes this a challenging proposition.