Jaded researchers cite unresolved alignment risks, limited external model review, and accelerating superintelligence development amid prospective public-market interest
An AI safety researcher’s public warning about the possibility of AI-driven human extinction has intensified scrutiny of the company as it faces growing speculation about a future public listing.
Evan Hubinger, who leads alignment science at Anthropic, has posted on X (on 8 September 2026) that he believes there is better than a 10% chance that AI could cause the deaths of all humans within the next 10 years. He said Anthropic is making a good-faith effort to manage the threat, but argued that the field has not solved the alignment problem for superintelligent systems, and is not plainly on a path to doing so.
Hubinger was replying to Jacob Coxon, an AI researcher who said he had left Anthropic after working for three years across Anthropic and OpenAI. Coxon said the major AI labs are pursuing self-improving superintelligence at a pace he considers dangerously reckless. He argued that people building frontier systems privately acknowledge the possibility of catastrophic harm more readily than they do in public.
Samuel Marks, who works on scalable oversight at Anthropic, subsequently joined the discussion and said concern about advanced-AI risks tends to increase among more senior-level employees.
Disparities between actions and promises
The exchange has reached Wall Street media as investors and commentators consider what a potential Anthropic initial public offering (IPO) could mean for the broader technology market. On CNBC’s Mad Money, Jim Cramer said the stock launch could prompt investors to shift money out of other investments to participate. The following morning, however, Cramer focused on Hubinger’s stated probability estimate, comparing a 10% fatality risk to the kind of risk a patient would likely reject before an elective operation. He said he was surprised Hubinger is still at the company after making the statement.
Hubinger’s assessment is broadly consistent with warnings made by other leading AI figures:
- Anthropic Chief Executive Dario Amodei had previously put the chance of catastrophic AI outcomes at approximately 25%
- Geoffrey Hinton has cited a range of 10% to 20%
- The comments are notable because Hubinger works directly on the technical challenge of ensuring that highly capable AI systems remain aligned with human intent and interests
The debate comes amid broader concern over whether AI developers are providing sufficient outside scrutiny of their most advanced models. The Financial Times recently reported that Anthropic did not provide its latest model to the United Kingdom’s AI Safety Institute. Neil Lawrence, a University of Cambridge machine-learning professor, told the BBC that he found the report credible and linked it to a geopolitical environment in which the United States increasingly frames AI development as a strategic competition with China.
Neither Anthropic nor OpenAI has indicated that it will halt work on more capable systems. Separately, an open letter signed by 1,300 employees of AI firms has urged the US government to help slow and more deliberately manage the development of fronter automated AI systems.