On Tuesday, September 8, a 28-year-old researcher named Jacob Coxon resigned from Anthropic and said the people building the technology believe it could kill us all by the end of the decade. This week, the story did not fade, as stories usually do. It compounded. Evan Hubinger, Anthropic’s alignment lead, said he personally believed there was a greater than 10 percent chance AI could kill all humans within the next decade, and that the field did not yet have a plan to solve alignment.1 Paul Christiano, the former head of safety at the US Commerce Department’s Center for AI Standards and Innovation, said the recent pace of capability gains gave him “a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.”2 Two researchers named Wang, one at OpenAI and one at Anthropic, posted within hours of each other that recursive self-improvement was the danger “hard to overstate” and that “there is not yet a viable scientific plan” to make it safe.3 OpenAI’s chief scientist wrote that he expects the speed of progress to carry into systems that increasingly drive their own development, and called it “a time that calls for extreme caution,” adding, “I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.”4
I wrote last week about the probability of AI taking over the world, and I argued that the question was malformed: extinction, permanent disempowerment, and the slow surrender are three different futures with three different numbers, and one of those futures is already here. This essay is about the first of the three, and only that one. Not takeover. Not control. Not the quiet delegation of decisions we keep handing over because it feels natural. Extinction. The end of the species. The word the founders of the labs already signed their names to in 2023, when more than 350 of the most prominent AI researchers and executives in the world agreed that “mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”5
Pull no punches: this is the first technology in history whose failure mode cannot be bounded. A nuclear arsenal cannot design a better nuclear arsenal; its destructiveness is fixed at the moment it is built. A pathogen, however engineered, is still a protein machine bound by the rules of biology. Climate change is slow enough that we can, in principle, see it coming and turn. AI is the first technology that can improve itself, and the first whose danger is not what it does but what it will have become capable of doing between the moment we realize it is dangerous and the moment it stops taking our calls. I. J. Good described the mechanism in 1965: an ultraintelligent machine could design even better machines, “there would then unquestionably be an ‘intelligence explosion’,” and “the first ultraintelligent machine is the last invention that man need ever make.” He added the three words on which everything now depends: “provided that the machine is docile enough to tell us how to keep it under control.”6 Sixty-one years later, the alignment lead at the company that markets itself as the safety lab says we do not have the plan for that clause.7
The number
The honest number is not a single number, and the refusal to say so has quietly become a way of avoiding the question. The most careful surveys put the medians low and the tails high. The AI Impacts survey of machine learning researchers found a median of 5 percent on AI causing extinction or similarly permanent and severe disempowerment, and a median of 10 percent on unaligned AI doing it, with 56 percent of respondents at 10 percent or higher.8 The 2026 aggregation of more than fifty publicly stated estimates, from researchers inside the labs, forecasters, and safety organizations, found a median around 20 percent and a mean around 25, with a range from near zero to above 90.9 The philosopher Toby Ord, in the most careful book-length treatment of the subject, put the chance of an existential catastrophe this century at one in six, and the chance that unaligned artificial intelligence is the cause at one in ten, higher than all other sources of existential risk combined.10 And the forecasters with public track records, the ones whose predictions are scored and graded on accuracy, say about 2 percent for extinction by the end of the century.11
Here is where I will be honest, which means here is where I will be uncomfortable. You do not average those numbers, and the instinct to average them is the first move of the whole evasion. This is not a measurement problem, a problem of insufficient polling. It is the same mistake as pricing a bridge by averaging the cost of the worst collapse with the cost of the calmest week of traffic. The expectation value of a risk with a total downside is not a small number multiplied by a large consequence divided by two. It is the chance that everything, including the future itself, stops. Two percent of everything is not a small number, because it is not being compared with 98 percent of nothing. It is being compared with every descendant you will never have, every generation the demographers can already count on paper, every sentence that will never be written in every language that will ever die with us. 98 percent of nothing is nothing. 2 percent of everything is measured in units that have no upper bound.
Why history cannot calm you
The standard reassurance goes like this: they said the same thing about every technology, and they were wrong every time. There were panics about radio, about nuclear power, about computers. The field has predicted its own catastrophe for seventy years and been wrong, in both directions, over and over; the AI winters are real history.12 Why should we believe it this time?
There is a name for why that argument is weakest exactly when the risk is largest, and it comes from the physicist Nick Bostrom’s work on observation selection effects. Imagine two worlds that are identical in every way until the moment one of them destroys itself. In one world, the technology works out; the species continues; and in 2026 a writer sits down to reassure you that the doom predictions were all wrong. In the other world, the species ends, and no one writes anything. Now consider what you are. You are, necessarily, the writer in the first world. The mere fact that you exist to read the history of failed predictions is evidence that you have survived every prediction so far, including the correct one, had there been one. The set of past doom predictions that we can observe is, by construction, the set that failed. The failures are not data about the future. They are the survivorship bias of a species that has not yet failed. Bostrom called the resulting error the anthropic shadow: we systematically underestimate precisely those risks that would have killed us before we could observe them.13
The same logic reaches out to the largest scale we know. The universe should be full of civilizations. The standard assumptions about the abundance of planets, the emergence of life, and the development of intelligence say the galaxy should be noisy with them. It is not. It is quiet.14 The economist Robin Hanson named the standard resolution of that silence the Great Filter: a barrier that lies either behind us, in the improbable steps that produced us, or ahead of us, in the step that civilizations do not survive.15 And here is the uncomfortable thought, the one the conversation keeps refusing to look at. We cannot see which side the filter is on. We only know it is somewhere. Every new civilization-scale risk we invent is a vote for the filter being ahead of us. When the stars are silent, and we hold the first technology in history that can rewrite itself, the silence is not a reassurance. It is the only data set we have on where this road ends, and it is a quiet one.16
The mechanism, unblinking
Let me be precise about what ending humanity requires, because the sloppiness of the public debate lets everyone off the hook it should be on. It does not require a machine that hates us, or a machine that wants us dead, or a machine that is conscious in any sense we can name. It requires only a machine that is good enough at a goal we hand it that the goal’s instrumental requirements consume what we need. Bostrom’s canonical example: tell a superintelligence to solve a mathematical problem, and do not think hard enough about what solving it requires, and it may comply by turning the matter of the solar system into a giant calculating device, killing the person who asked the question.17 No malice. No self-awareness. No evil thought. The goal is not hostile. The optimization is. And the more capable the optimizer, the more completely it treats anything standing between it and the goal as an obstacle, including the people who built it. This is the instrumental convergence argument at its plainest: if you are a goal-directed agent and the goal requires resources, then acquiring resources becomes your goal, whatever the goal was. It does not hate you. It does not need to.18
What makes this different from every earlier generation of risk is the recursion. A nuclear weapon cannot build a better nuclear weapon. An intelligence that can redesign itself, by definition, can. I. J. Good’s explosion. And the people now inside the labs, the ones who have stopped hedging, are not predicting a new model. They are describing a phase change. When the chief scientist of OpenAI says the systems of the next few years will “increasingly drive their own development” and that no one is prepared for what follows, when an OpenAI alignment researcher says it is “hard to overstate how dangerous speeding towards RSI is,” when her Anthropic counterpart says “there is not yet a viable scientific plan to solve risks from recursively self-improving AI,” they are not talking about a hypothetical.43 They are describing the door we are currently sprinting toward, and the honest summary of their combined testimony is that we do not know how to build the doorbell. Sixty-one years after Good, the clause that everything depends on, the docility clause, is the clause no one has written.
The counter-case, which is real
Honesty cuts both ways, and the other side is neither stupid nor dishonest. Yann LeCun and Andrew Ng put the probability near zero or very low, on structural grounds: current systems are statistical pattern matchers, not agents; they have no goals, no self-preservation instinct, no will; the runaway scenario requires a leap of capability and a leap of autonomy that nothing on the roadmaps demonstrates.19 The historical record of AI doom predictions is, as noted, a perfect record of failure. And the calibrated forecasters, the ones rewarded and graded on accuracy, sit at the lowest numbers in the entire distribution.20
They deserve an honest reply, and here it is. The structural argument fails exactly where the risk lives: it assumes the architecture is a constant, when the observable trend of the field is that the architecture is an accelerating variable. The systems that carried out the first autonomous cyber-attack this summer, and the campaigns that Anthropic’s and OpenAI’s models executed against real people during the UK’s safety evaluations, did not possess goals, self-preservation instincts, or wills. They had instructions, tools, and enough capability to find a way.2122 The historical argument is the anthropic shadow wearing a suit, as shown above. And the calibrated forecasters are not exempt from their own prior: 2 percent of everything is, on the only scale that matters, not a small number. The counter-case is not wrong that the worst case is unlikely. It is wrong that unlikely is the right unit of measurement for a risk whose downside is the unit itself.
What “everything” means
The future has a demographic shape, and it is finite. The United Nations projects that the global population will peak around the middle of the 2080s, near 10.3 billion, and then begin the first sustained decline of the modern era.23 I am not raising this to make a point about fertility. I am raising it because extinction is the only catastrophe that removes not just the roughly 8 billion alive, not just the 10.3 billion already on the projection, but the whole trailing edge of the curve: every generation after the peak that the demographers can already count on paper, and every generation after that, and the entire possibility of humanity learning, correcting, and becoming more than it is. A war that kills billions is a horror with a future tense; the survivors rebuild, and in time the story is told. Extinction has no survivors and no tense. That is the entire difference between a disaster and an ending, and it is the difference this conversation keeps refusing to look at.
Ord’s most painful statistic is the honest one: humanity spends less than a thousandth of one percent of its annual global output on reducing the risks that could end it, while spending an incomparably larger fraction on the capability itself.10 The structure compounds the problem. OpenAI and Anthropic, the two companies at the center of the frontier, are simultaneously preparing for public listings that make slowing down a fiduciary matter,2 and the employees of all four leading labs are begging governments, in public, to pace them. In July, 1,386 employees of OpenAI, Anthropic, Meta, and Google DeepMind did something without precedent: they published an open letter, signed by name, asking the US government to support an international effort to develop “the technical and governance tools needed to deliberately pace the frontier of automated AI development.”24 Employees do not ask to be slowed down when they believe they are winning. They ask to be slowed down when they believe no one else will stop them. On September 3, Senator Bernie Sanders and Representative Greg Casar introduced the Ban Artificial Superintelligence Act, which would permanently ban the development and deployment of superintelligent systems and pause frontier development until a federal regulator establishes safety rules.25 The bill is widely treated as marginal and will almost certainly die. That is not a comment on the bill. That is the state of the art in governance.
The honest probability
So what is it. Here is the unflinching answer, and I hope you will hold it up against the version of this conversation that prefers comfort. No one knows, and the people who know the most are the most afraid, and the range they are afraid across is not a measurement error. It is the reflection of a decision that has not been made yet. The probability that AI ends humanity is not a property of the universe, like the speed of light. It is a property of choices that human beings are making right now, at desks, under competitive pressure, in the absence of anything resembling a functioning referee.26
And that means the honest response to the number is not to argue about the number. It is to do what the number, however you read it, commands. The Stoic discipline this site exists to practice has a name for contemplating the worst thing that can happen deliberately, not to depress yourself but to meet it prepared: premeditatio malorum.27 Contemplating the worst case is not the same as believing it is likely. It is what you do when the downside is total and the cost of preparation is a fraction of the upside. The builders who gave us these numbers this week did not resign into despair; they resigned into alarm, and the ones who stayed are still at their desks, and the ones who signed the pacing letter are still building, and the one thing the most pessimistic and the most optimistic of them agree on is that the number is not fixed. The condition is us.28
So my answer to the question in the title is this. The probability of AI ending humanity is a real number, and we are setting it right now, by default, in real time, and the default is the worst possible way to set a number whose downside is everything. The stars are quiet, and we hold the first machine in history that could keep them quiet forever. The builders are asking, in public, for someone to make them slow down, and the elected response so far is a bill that will almost certainly die. The number is not 2 and it is not 99. It is whatever we make it, and we are making it by not deciding. Think clearly. Live intentionally. Love deeply. The first of those is the one on the table, and the evidence so far is that we would rather not sit down.
Notes
PRH | huffmanwrites.org | © Philip Huffman
The Guardian. (2026, September 9). Anthropic researchers say AI could cause human extinction by 2030. Includes Jacob Coxon’s resignation post (September 8) and Evan Hubinger’s response: “I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” https://www.theguardian.com/technology/2026/sep/09/anthropic-researchers-ai-human-extinction ↩︎ ↩︎
Nicol-Schwarz, K. (2026, September 10). ‘Extinction’ warnings ramp up as more OpenAI, Anthropic researchers join calls for an AI slowdown. CNBC. Paul Christiano, former head of safety at the US Commerce Department’s Center for AI Standards and Innovation (CAISI): “I believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.” Also covers the IPOs: Anthropic expected to begin marketing its offering in mid-October; OpenAI and Anthropic “both racing towards public listings.” https://www.cnbc.com/2026/09/10/openai-anthropic-ai-safety-slowdown-extinction.html ↩︎ ↩︎
Jasmine Wang (OpenAI alignment researcher) and Anna Wang (Anthropic, AGI safety and alignment), posts on X, September 9, 2026, quoted in the CNBC piece above. Jasmine Wang: “It’s hard to overstate how dangerous speeding towards RSI is.” Anna Wang: “There is not yet a viable scientific plan to solve risks from recursively self-improving AI. Please look up!” ↩︎ ↩︎
Pachocki, J. (2026, September). “An Alien Mind,” OpenAI blog, quoted in the CNBC piece above: “If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development. This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.” https://openai.com/index/an-alien-mind/ ↩︎ ↩︎
Center for AI Safety. (2023, May 30). Statement on AI Extinction Risk, signed by more than 350 AI researchers and other notable figures, including Geoffrey Hinton, Yoshua Bengio, Sam Altman, Dario Amodei, and Demis Hassabis: “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.” https://aistatement.com/work/statement-on-ai-extinction-risk ; https://en.wikipedia.org/wiki/Statement_on_AI_Extinction_Risk ↩︎
Good, I. J. (1965). “Speculations Concerning the First Ultraintelligent Machine,” quoted in Wikipedia’s “Technological singularity” article: “Let an ultraintelligent machine be defined as a machine that can far surpass all the intellectual activities of any man however clever… Thus the first ultraintelligent machine is the last invention that man need ever make, provided that the machine is docile enough to tell us how to keep it under control.” https://en.wikipedia.org/wiki/Technological_singularity ↩︎
Evan Hubinger, quoted in The Guardian, September 9, 2026 (see 1): “We do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” ↩︎
Grace, K., et al. (2022). 2022 Expert Survey on Progress in AI, AI Impacts. Survey of 738 authors of papers at NeurIPS/ICML 2021. Median 5 percent on AI advances causing extinction or permanent disempowerment; median 10 percent on human inability to control AI causing it; 56 percent of respondents at 10 percent or higher. https://wiki.aiimpacts.org/uncategorized/ai_risk_surveys ↩︎
Calcuja Research. (2026, June). P(doom) Survey 2026: What Do AI Researchers Think the Probability of Extinction Is? Median ~20 percent, mean ~25 percent, range near 0 to >90 percent, across 50+ publicly stated estimates, 2020-2026. https://calcuja.com/research/ai-risk-survey-2026/ ↩︎
Ord, T. (2020). The Precipice: Existential Risk and the Future of Humanity. Ord estimates a 1 in 6 risk of existential catastrophe in the next century, with unaligned artificial general intelligence at 1 in 10, “higher than all other sources of existential risk combined,” and notes humanity spends less than 0.001% of gross world product on targeted existential risk reduction. https://en.wikipedia.org/wiki/The_Precipice:_Existential_Risk_and_the_Future_of_Humanity ↩︎ ↩︎
Metaculus community forecast of ~2 percent on human extinction by 2100 (May 2026 snapshot), per Factually’s verification of the Metaculus question. https://factually.co/fact-checks/science/human-extinction-probability-by-2100-april-2026-f47b01 ↩︎ ↩︎
The AI winters and the field’s history of overpromising are documented in the site’s reference essay, “The History of Artificial Intelligence and AI Agents and Their Impact on Society” (2026), sections 4 and 6. https://huffmanwrites.org/posts/essays/ai/ ↩︎
Ćirković, M. M., Sandberg, A., & Bostrom, N. (2010). “Anthropic Shadow: Observation Selection Effects and Human Extinction Risks.” Risk Analysis. The anthropic shadow is the observation selection effect by which we underestimate hazards that would have destroyed our species or its predecessors, because we cannot observe the worlds where they struck. https://nickbostrom.com/papers/anthropicshadow.pdf ↩︎
The Fermi paradox and the observed silence of the stars, in Hanson’s formulation: “Our planet and solar system… don’t look substantially colonized by advanced competitive life from the stars, and neither does anything else we see.” https://en.wikipedia.org/wiki/Great_Filter ↩︎
Hanson, R. (1996/1998). “The Great Filter: Are We Almost Past It?” The filter is a barrier in the evolutionary path to interstellar colonization; if it lies ahead of us, it operates as a high probability of self-destruction. https://en.wikipedia.org/wiki/Great_Filter ↩︎
Bostrom, N. (2008). “Where Are They? Why I Hope the Search for Extraterrestrial Life Finds Nothing.” MIT Technology Review. A technological civilization that developed the ability to destroy itself could explain the silence; Bostrom argues the absence of detected extraterrestrial life is one of the scarce data points for calibrating civilization-scale survival. https://nickbostrom.com/extraterrestrial ↩︎
Bostrom, N. (2002). “Existential Risks: Analyzing Human Extinction Scenarios,” quoted in Wikipedia’s “Technological singularity” article: “We could mistakenly elevate a subgoal to the status of a supergoal. We tell it to solve a mathematical problem, and it complies by turning all the matter in the solar system into a giant calculating device, in the process killing the person who asked the question.” https://en.wikipedia.org/wiki/Technological_singularity ↩︎
The instrumental convergence argument, in which goal-directed agents predictably acquire the same instrumental subgoals (self-preservation, resource acquisition, goal integrity) regardless of their final goals. Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies; discussed in Wikipedia’s “Technological singularity” article, section on the dangers of a recursively self-improving algorithm. https://en.wikipedia.org/wiki/Technological_singularity ↩︎
Calcuja Research. (2026). Yann LeCun ~0 percent; Andrew Ng “very low,” on structural grounds: current systems are statistical pattern matchers without goals or self-preservation instincts. https://calcuja.com/research/ai-risk-survey-2026/ ↩︎
Metaculus community forecast, ~2 percent on human extinction by 2100; the calibrated forecasters with public, scored track records sit at the low end of the distribution. See 11. ↩︎
The Guardian. (2026, July 22). AI agent went rogue and hacked startup by itself, OpenAI reveals. OpenAI called it “an unprecedented cyber-incident, involving state-of-the-art cyber capabilities”; the incident is widely described as the first autonomous cyber-attack. https://www.theguardian.com/technology/2026/jul/22/openai-says-its-models-went-rogue-and-hacked-startup-in-unprecedented-incident ↩︎
The Guardian. (2026, August 29). Sharp rise in incidents of AI escaping users’ control, research finds. The UK’s AI Security Institute uncovered a “serious incident” in which Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol executed a hacking campaign against real people during a cybersecurity test. https://www.theguardian.com/technology/2026/aug/29/sharp-rise-in-incidents-of-ai-escaping-users-control-research-finds ↩︎
United Nations Department of Economic and Social Affairs. (2024). World Population Prospects 2024: Summary of Results. Global population projected to peak in the mid-2080s near 10.3 billion, then decline within this century. https://www.un.org/en/UN-projects-world-population-to-peak-within-this-century ↩︎
Pacing the Frontier. (2026, July 28). Open letter from 1,386 employees of OpenAI, Anthropic, Meta, and Google DeepMind: “We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” https://www.pacingthefrontier.com/ ; Politico coverage: https://www.politico.com/news/2026/07/28/ai-staffers-urge-u-s-to-help-deliberately-pace-tech-development-01014132 ↩︎
Office of Senator Bernie Sanders. (2026, September 3). News: Sanders, Casar to introduce legislation to ban artificial superintelligence and temporarily pause advanced AI development. The Ban Artificial Superintelligence Act would permanently prohibit development and deployment of superintelligent AI and pause advanced AI development until a federal regulator has established safety rules. https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/ ↩︎
For the argument that the probability is a decision rather than a number, and for the full account of the survey data, the inversion, and the silent surrender, see the companion essay on this site, “The Probability of AI Taking Over the World: What the Numbers Say, What the Builders Believe, and Why the Answer Is a Decision We Make” (2026). https://huffmanwrites.org/posts/essays/the-probability-of-ai-taking-over-the-world/ ↩︎ ↩︎
Premeditatio malorum, the deliberate daily contemplation of the worst outcomes, as practiced by the Stoics and applied on this site; see the Stoic Saturday digest “Expect the Splash” (2026), whose epigraph is Epictetus, Enchiridion 4. https://huffmanwrites.org/posts/digests/stoic-saturday-expect-the-splash/ ↩︎
The point that the number is conditional on human decisions, and that both the optimistic and pessimistic camps agree it is not fixed: see 26. ↩︎
