The Probability of AI Ending Humanity: What the Builders Just Said, Why the Stars Are Quiet, and Why 2 Percent of Everything Is Not a Small Number
On Tuesday, September 8, a 28-year-old researcher named Jacob Coxon resigned from Anthropic and said the people building the technology believe it could kill us all by the end of the decade. This week, the story did not fade, as stories usually do. It compounded. Evan Hubinger, Anthropic’s alignment lead, said he personally believed there was a greater than 10 percent chance AI could kill all humans within the next decade, and that the field did not yet have a plan to solve alignment.1 Paul Christiano, the former head of safety at the US Commerce Department’s Center for AI Standards and Innovation, said the recent pace of capability gains gave him “a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.”2 Two researchers named Wang, one at OpenAI and one at Anthropic, posted within hours of each other that recursive self-improvement was the danger “hard to overstate” and that “there is not yet a viable scientific plan” to make it safe.3 OpenAI’s chief scientist wrote that he expects the speed of progress to carry into systems that increasingly drive their own development, and called it “a time that calls for extreme caution,” adding, “I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.”4 ...