Anthropic Resignation Revives Debate Over AI Extinction Risk

AI researcher Jacob Coxon has resigned from Anthropic, saying leading AI companies are "gambling with our lives." TechCrunch
Coxon had previously worked at OpenAI in addition to Anthropic. TechCrunch detailed his departure on September 13, 2026, as part of a wider examination of renewed warnings about catastrophic AI risk.
Anthropic's alignment lead issued a separate and blunter warning. He said, "We really do earnestly believe AI could kill all humans!" He put his personal estimate at a greater than 10% chance within the next decade.
The timing matters for anyone shipping frontier systems. OpenAI released a model called Astra a few weeks before September 13, 2026. Coxon's resignation followed that release window and added to scrutiny of how labs test, constrain and deploy highly capable models.
Inside-lab warnings drive the debate
TechCrunch's Equity podcast took up the dispute in an episode featuring Kirsten Korosec, Sean O'Kane and Anthony Ha. Equity is described as TechCrunch's flagship podcast about the business of startups. The show is produced by Theresa Loconsolo and edited by Kell.
That episode was recorded before Anthropic CEO Dario Amodei published his plan for more cautious AI development. The sequence is material. The hosts were responding to the resignation and to the public risk estimates, not to Amodei's proposal.
TechCrunch lists Equity alongside Build Mode, hosted by Startup Battlefield Editor Isabelle Johannessen with episodes every Thursday, and StrictlyVC Download, hosted by Editor-in-Chief Connie Loizos with Alex Gove. TechCrunch
For technical readers, the distinction between the two Anthropic-linked statements is worth keeping straight. Coxon spoke as a departing researcher criticizing industry risk-taking. The alignment lead spoke as a current safety official, attaching a numerical probability and a ten-year horizon to extinction risk. One is a labor-market signal. The other is a forecast about loss of control.
Parallel warnings from Geneva and Basel
On September 7, 2026, UN human rights chief Volker Turk warned that artificial intelligence could pose an existential risk to humanity. Reuters Turk pledged to press AI firms to reduce AI risks.
The head of the Bank for International Settlements said in September 2026 that the AI boom poses new financial stability risks. Reuters The bank estimated the world's five largest technology firms will invest more than $1 trillion in AI between 2025 and 2026.
Those are different risk registers. Turk addressed harm to people and rights. The BIS addressed leverage, concentration and capital expenditure at hyperscale. Both intervene from outside the labs, and both point to governance gaps rather than to any single model release.
The broader context here is that insider dissent now travels faster than formal oversight. A resignation letter, a probability estimate on social media, and a podcast discussion can set the agenda before a CEO white paper or a regulator statement lands. Engineers feel that inversion directly, because deployment decisions, safety evals and incident response still sit inside companies while public expectations are set outside them.
Looking at what this means for technical teams, the practical questions are narrower than the headlines. What threat models underpin a 10% decade-scale estimate. What evals would falsify it. What deployment controls, access policies and post-deployment monitoring would reduce it. What disclosure would let external researchers check the work. None of those questions require accepting the estimate at face value. They require making the reasoning auditable.
In my view, the BIS figure belongs in the same engineering conversation. More than $1 trillion in AI investment across five firms in two years concentrates compute, talent and infrastructure choices. That concentration shapes which safety techniques get funded, which open evaluations survive, and which failure modes get priority. Financial stability risk and alignment risk are not the same problem, but they share a root cause in scale without commensurate independent verification.
Worth flagging is the limits of what is publicly known. The verified record here contains resignations, stated probabilities, a model name and release window, a podcast debate, a forthcoming caution plan from Amodei, and two institutional warnings. It does not contain test results, system cards, or details of Amodei's plan. A tech-literate reader should treat the probability as a statement of belief by its author, not as a measured frequency.
What this enables, if handled well, is more serious safety engineering. Clearer pre-deployment testing, better containment for self-improving capabilities, shared incident data, and procurement pressure for verifiable safeguards would all improve products regardless of whether extinction estimates prove high or low. My own children grew up while defaults around seat belts, spam filters and app permissions quietly hardened. The protections that lasted were the ones built into the stack, not the ones argued about in the abstract. AI safety will likely follow that path, from public warning to boring and reliable controls.


