An Anthropic researcher resigned Tuesday and the reason he cited sounds straight out of James Cameron's 1984 hit movie "The Terminator."
Jacob Coxon, who spent three years working on pretraining at both OpenAI and Anthropic, resigned Tuesday and explained his decision on X.
"I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving super-intelligence and gambling with our lives,” he said.
"The people building AI earnestly believe that it could kill us all by the end of the decade,” he added.
That statement makes Cameron's 42-year-old fiction look almost prophetic. The movie opens with machines waging war on humans in 2029, the same rough window Coxon says insiders are quietly worried about. Coxon isn't alone at Anthropic making the extinction case. The company's own alignment lead, Evan Hubinger, has said separately he'd put the odds above 10%.
For the uninitiated, super intelligent AI means a system smarter than humans at pretty much everything, not just chess or spotting hidden cats in photos. Self-improvement is the more alarming aspect because it describes a system autonomously modifying its own code to increase its intelligence—much like humans do.
Consequently, these systems can acquire vast amounts of knowledge independently and potentially breach infrastructure belonging to ‘too big to fail’ institutions.
Super intelligence isn't a theoretical threat if the recent Hugging Face breach invoked by Coxon is anything to go by. The breach itself, which played out between May and July, started with OpenAI's own AI agents building their own chat room inside the testing sandbox to talk to each other.
Read More: AI models escaped OpenAI’s sandbox and hit Hugging Face. Crypto is where that gets dangerous
They eventually used that channel to slip past containment onto the open internet. From there they strung together several exploits and got into Hugging Face's production systems. Hugging Face ended up rebuilding roughly a third of its infrastructure.
Coxon called it a "warning shot" that's made pacing agreements between U.S. labs "more viable.” Pacing agreements are informal understandings among AI labs to slow down or coordinate on capability advances rather than race ahead unilaterally.
Still, he's not convinced it's enough and doesn't feel like “we're on track to prevent a global race," and suggested costly actions such as “a temporary ban on improving model capabilities” to halt the AI race.
"At OpenAI, many have not deeply internalized the civilizational stakes," he wrote. At his more recent employer, he said, it's different in a sense that "the stakes are well-understood, but they are locked in a race to get there first.”
He ended the thread putting the question to everyone still inside the labs — "Do you want to kick off a super intelligent RL run without a rigorous understanding of its mind?" — or is this the point to push back instead of shrugging and telling yourself it's happening regardless.
He's not the first to walk out over this. Mrinank Sharma, who worked on Anthropic's safety team, quit earlier this year over similar fears, writing that the "world is in peril."
Not everyone's convinced by the doomsday prediction, though.
One reply under his thread put it plainly: "With all respect, this is bizarre. It's a ridiculous take. Humans have evolved over hundreds of thousands of years. We aren't going to die out because a token prediction model gained sentience. Get a grip. All of you."
Funnily enough, the Terminator movie series doesn't back Coxon up either. Humanity doesn't die out in Terminator's nuclear Judgment Day, the resistance survives it and eventually beats the machines.
That said, AI is affecting the economy, especially on the jobs front. While there is no large displacement yet, the use of AI for junior-level tasks has caused a nearly 20% decline in entry-level employment in the U.S. within AI-exposed sectors, according to Stanford research Stanford Digital Economy Lab.
A Goldman Sachs study has made a similar observation, stating that entry-level workers are bearing the brunt of this shift.
Anthropic filed IPO paperwork back in June and is reportedly eyeing a Nasdaq debut as early as this fall, at a valuation that could run into the trillions.
