deniz.in

Markets

Weather

Loading weather

· via The Verge

Anthropic safety lead puts odds of AI killing humanity above 10% as colleague resigns

Anthropic researcher Jacob Coxon quit over what he called a reckless race toward superintelligence; safety lead Evan Hubinger replied that he puts the odds of AI killing everyone above 10 percent and Anthropic lacks a safety plan.

Anthropic safety lead puts odds of AI killing humanity above 10% as colleague resigns

A resignation over the race to superintelligence

On September 9, Jacob Coxon, a researcher who has trained AI systems at Anthropic and previously at OpenAI, announced his departure in a post on X. According to The Verge, Coxon said he quit over what he sees as an insufficient commitment to safety, accusing both labs of "racing straight to self-improving superintelligence and gambling with our lives" — despite the fact that "the people building AI earnestly believe that it could kill us all by the end of the decade."

In his telling, the companies are "locked in a race" to build advanced systems ahead of their rivals and are pushing forward "despite the risk." The core worry is recursive self-improvement: a runaway loop in which AI systems keep upgrading themselves beyond human control. The Verge notes that this outcome has not materialized, but companies are actively pursuing it, and much of the code behind today's AI is already written with AI assistance.

Anthropic's own safety lead puts the risk above 10 percent

Within hours, Evan Hubinger, who leads one of Anthropic's AI safety teams, replied directly — and largely agreed. He said he worries about self-improving AI and that it "is happening faster than we thought," and he endorsed Coxon's characterization of the situation.

"We really do earnestly believe AI could kill all humans," Hubinger wrote, adding that he personally estimates the probability at greater than one in 10 "within the next decade."

The most consequential line, as reported by The Verge, was Hubinger's admission that Anthropic does "not yet have a plan" for ensuring advanced AI remains safe and aligned with human values — and that the company is "not clearly on track to" develop one.

A lab built on safety worries now faces its own

Anthropic was founded by former OpenAI employees who left that company over safety concerns, which gives Coxon's exit an uncomfortable symmetry. In recent years, multiple researchers have also left OpenAI while citing safety as their motivation, and The Verge describes Coxon's departure as one of the most high-profile exits from Anthropic to date.

The exchange also lands at a delicate moment for the industry. According to The Verge, the major labs are preparing for anticipated IPOs while managing the fallout from numerous incidents involving rogue AI agents, alongside high-profile warnings about how difficult frontier models are to monitor. The story spread quickly beyond X: Politico covered it under a headline borrowing Coxon's "gambling with our lives" phrase, and the article reached the front page of Hacker News.

Why it matters

When the people responsible for making frontier AI safe state publicly, in their own words, that they see a better-than-10-percent chance of it killing everyone within a decade — and that no workable plan for preventing that outcome exists yet — the usual industry reassurances become much harder to sustain. This is not an outside critic or a pundit; it is a serving safety lead at one of the world's leading AI labs, responding to a resignation from inside the same building.

That it happened at Anthropic, a company whose founding story is safety, suggests that competitive pressure applies even to labs with the strongest safety branding. If a safety researcher concludes that quitting is the strongest statement he can make, and his colleague responds by confirming the risk estimate while acknowledging the missing plan, the open question is whether market dynamics leave any lab room to slow down — and who, if anyone, is positioned to make them.

  • #ai-safety
  • #anthropic
  • #openai
  • #superintelligence
  • #existential-risk

Related posts