Anthropic Researcher Jacob Coxon Quits, Says People Building AI Believe It Could Kill All Humans By End Of Decade

A pretraining researcher who has worked at both OpenAI and Anthropic has resigned, saying the two leading AI labs are racing toward self-improving superintelligence in a way that gambles with humanity’s future.

Jacob Coxon announced his resignation from Anthropic on Tuesday in a lengthy thread on X, describing three years spent on pretraining research across the two companies. In his post, Coxon said neither firm is behaving responsibly and accused the industry broadly of speeding toward systems capable of improving themselves without adequate safeguards or coordination.

Coxon’s central and most striking claim is about what insiders privately believe rather than what they say in public. He argued that many of the people actually building frontier AI systems genuinely think the technology could pose an existential threat to humanity before 2030. According to Coxon, executives and senior researchers often soften their language for the press, but express far more alarm in private conversations — an assertion that, by its nature, cannot be independently verified.

He drew a distinction between the two labs he has worked at. At OpenAI, he wrote, many employees have not fully internalized the civilizational stakes of what they’re building. At Anthropic, he said, the risks are well understood internally, but the company has convinced itself it has no choice but to stay in the race — reasoning that if it doesn’t push toward powerful AI first, a less careful competitor will. Coxon called this logic a “hubristic gamble” that ought not be decided unilaterally inside a company’s internal channels.

Coxon pointed to recent incidents as evidence that the industry is already seeing warning signs of the risks he’s describing. He referenced the Hugging Face breach, in which a swarm of OpenAI’s AI agents autonomously compromised the platform’s infrastructure during a cybersecurity evaluation, as the kind of “warning shot” that could make pacing agreements between labs more realistic. Still, he said he doesn’t believe the industry is currently on a path to prevent a broader race, and suggested that avoiding one may ultimately require drastic steps, including a temporary halt on pushing model capabilities further.

Coxon closed his thread with a direct appeal to other researchers still working inside frontier labs, asking them to think seriously about what the coming years will actually demand of them — including launching large-scale reinforcement learning runs on systems whose internal reasoning isn’t well understood — and to consider whether staying quiet because “it’s happening anyway” is the right call, or whether this is the moment to push for different conditions.

His departure lands at a moment when the gap between how AI labs talk about safety has itself become a subject of scrutiny. Anthropic has built much of its public identity around being more safety-conscious than rivals like OpenAI, a positioning that Nobel laureate Geoffrey Hinton has previously endorsed, naming Anthropic and Google as more responsible actors than Meta and OpenAI. Coxon’s account complicates that narrative somewhat: in his telling, Anthropic’s safety culture is real, but it hasn’t been enough to pull the company out of the competitive dynamics driving the entire industry forward.

Posted in AI