Close
newsletters Newsletters
X Instagram Youtube

Anthropic researcher sees over 10% chance AI could wipe out humanity within decade

An AI-generated illustration visualizes concerns over superintelligence, alignment and the potential catastrophic risks posed by advanced AI systems. (AI-generated illustration/Türkiye Today)
Photo
BigPhoto
An AI-generated illustration visualizes concerns over superintelligence, alignment and the potential catastrophic risks posed by advanced AI systems. (AI-generated illustration/Türkiye Today)
September 09, 2026 01:42 PM GMT+03:00

Anthropic Alignment Science lead Evan Hubinger said he believes there is a greater than 10% chance that artificial intelligence could kill all humans within the next decade.

He backed concerns raised by researcher Jacob Coxon, who announced his resignation from Anthropic while accusing leading AI companies of racing toward self-improving superintelligence without an adequate safety plan.

Hubinger stressed that the estimate reflected his personal view rather than Anthropic's assessment of its current systems.

He said the company was trying to address the danger, but argued that researchers still lacked a clear solution to the alignment problem for superintelligence, meaning the challenge of ensuring extremely capable AI systems continue to behave in ways consistent with intended goals and safety constraints.

TWEET

"Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to," Hubinger wrote.

TWEET

Coxon leaves Anthropic warning of race toward superintelligence

Coxon, an AI researcher who said he spent the past three years working on pretraining research at OpenAI and Anthropic, announced his departure from Anthropic on Sept. 9, saying he believed neither company was acting responsibly.

He argued that the companies were pushing ahead toward "self-improving superintelligence," referring to AI systems powerful enough to help accelerate the development of increasingly capable AI, while taking risks that could affect humanity as a whole.

Coxon said rapidly advancing systems could eventually become superhuman across areas including hacking and scientific research while gaining access to greater power and resources.

TWEET

"The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger," he wrote.

Coxon also addressed the apparent contradiction between believing such systems could pose catastrophic risks and continuing to develop them. He argued that many people at OpenAI had not fully taken in what he described as the civilizational stakes, while researchers at Anthropic understood the risks but believed they had to stay in the race because competitors might act less responsibly.

He described that approach as a gamble and called for stronger coordination among AI laboratories, arguing that measures could ultimately include temporarily restricting further improvements in model capabilities.

TWEET

Hubinger distinguishes future superintelligence from today's models

After his initial post, Hubinger clarified that his concern centered on future superintelligence rather than Anthropic's present systems.

He said he considered the risk from current models low, pointing to Anthropic's latest Risk Report, while identifying recursive self-improvement as the central concern.

In this context, recursive self-improvement refers to increasingly capable AI contributing to research and development that can in turn speed up the creation of still more capable systems.

Anthropic's August 2026 Risk Report similarly rates the current risk from automated AI research and development as "low," although it says the company is less confident in that assessment than in earlier reports because some evaluations no longer capture increases in model capabilities and because Anthropic is seeing early signs of acceleration.

The report also says Anthropic's internal AI research and development is already significantly faster with AI assistance, although the company does not believe the improvement has yet doubled its overall rate of progress.

At the same time, Anthropic acknowledges that its existing safeguards would not be enough to keep risks low in a scenario involving highly automated or dramatically accelerated research and development.

Anthropic report assesses present autonomy risks as low

The distinction between current and future systems also appears elsewhere in the company's report. Anthropic says autonomy-related threats from its present models are currently low and argues that their relatively weak covert capabilities limit their ability to bring about catastrophic harm.

However, the report says highly capable AI could rapidly accelerate research across technical fields, potentially disrupting balances of power if controlled by humans or producing catastrophic harm if such capabilities were combined with dangerous autonomous goals.

Coxon's resignation and Hubinger's comments therefore draw a distinction between assessments of today's AI systems and concerns over where continued capability development could lead.

While Anthropic's formal report continues to characterize present autonomy and automated R&D risks as low, both researchers focused their warnings on the prospect of much more capable systems emerging as AI-driven research speeds up.

September 09, 2026 01:42 PM GMT+03:00
More From Türkiye Today