An Anthropic researcher has resigned over AI safety concerns, warning that labs are racing toward self-improving superintelligence, sparking a heated online debate

A researcher who spent years working at OpenAI and Anthropic has walked away from the AI race, claiming the industry is moving toward a dangerous point of no return. His blunt warning has now gone viral online.
Jacob Coxon said he resigned from Anthropic after spending the past three years conducting pretraining research at both OpenAI and Anthropic. In a post that has since gone viral on X, Coxon alleged that neither company is acting ‘responsibly’ enough as the AI industry races toward increasingly autonomous and potentially self-improving systems.
Coxon warned that the industry is “racing straight to self-improving superintelligence” while, in his view, gambling with humanity’s future. He argued that the rapid development of AI requires much stronger safety measures and greater consideration of the potential consequences before systems become significantly more capable.
Elaborating on his concerns in the comments, Coxon urged people not to underestimate what advanced AI could eventually do. He argued that future systems could become “superhuman,” potentially capable of hacking systems, transforming entire fields in a short period and gaining access to real-world resources and power.
“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately. No other human activity poses this level of danger.”
Coxon also addressed the frequently raised question of why researchers would continue developing AI if they genuinely believe it could pose an existential threat.
According to him, the situation differs between major AI companies. He claimed that while some researchers at OpenAI may not have fully internalised what he described as the “civilizational stakes,” Anthropic researchers understand the risks but feel compelled to compete because of fears that another company or country could move ahead.
He described entering an AI “endgame” as a dangerous gamble and questioned whether decisions with potentially global consequences should be made within private companies.
“Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.”
Despite his concerns, Coxon said he remains optimistic that international coordination could still help prevent an uncontrolled AI race. He pointed to recent incidents, including what he referred to as the Hugging Face attack, as potential “warning shots” that could encourage AI laboratories in the US to agree on limits around development.
Coxon suggested that preventing a global race could eventually require significant measures, including temporarily restricting improvements in model capabilities.
He also directly appealed to AI researchers to consider what the next few years of development could look like and whether they are comfortable launching increasingly powerful reinforcement-learning systems without having a rigorous understanding of how those systems work internally.
“If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” – or take this moment to call for different conditions?”
Coxon’s resignation has sparked a debate online, with users divided over whether his warnings represent a serious and necessary intervention or an exaggerated view of AI’s potential risks.
“Could someone explain to me how AI could “kill us all” without using a nuke or some type of biological means as an example? I’m not even saying I don’t believe it could, I just genuinely don’t understand how AI can literally kill us all. I mean I understand the social and economic impact to a degree but I’m talking about this doomsday scenario people keep warning us about but can’t quite seem to be specific about,” said one user.
Another commenter dismissed the warning altogether, writing, “With all respect, this is bizarre. It’s a ridiculous take. Humans have evolved over hundreds of thousands of years. We aren’t going to die out because a token prediction model gained sentience. Get a grip. All of you.”

