Anthropic’s Alignment Lead Says the Company Has No Plan to Control Superintelligence

A 27-year-old pretraining researcher resigned from Anthropic on Tuesday, and the internet spent the next two days arguing about whether artificial intelligence is going to kill everyone. That is the wrong argument, and it is also the one the industry finds most comfortable having. Buried underneath the extinction headlines is a sentence from Evan Hubinger, who runs Anthropic’s Alignment Science team, that no communications department would ever have written: the company does not yet have a plan to solve alignment for superintelligence, and is not clearly on track to get one. Nobody at Anthropic has disputed him. It is the most checkable statement anyone made this week, and nearly every writeup filed it below the number.

What Coxon Actually Said

Jacob Coxon spent roughly three years doing pretraining research at OpenAI and then Anthropic before quitting this week. NBC News reported that he told colleagues on Slack that superintelligent systems carry a risk of causing human extinction, then posted a seven-part thread publicly. “Neither company is acting responsibly,” he wrote. “They are racing straight to self-improving superintelligence and gambling with our lives.” He asked for government intervention or a coordinated industry slowdown, including a temporary ban on improving raw model capability, which is a considerably more concrete ask than the genre usually produces.

A blank floor-to-ceiling whiteboard with faint erased marks in a darkened glass-walled conference room, an empty chair turned away from it
The gap between what the safety leads say in public and what the company has actually written down is the whole story this week.

The Confirmation Is the Story

Departures like this normally end with a company statement about differing views and a news cycle that dies by Friday. This one did not, because two people still on Anthropic’s payroll publicly agreed with him. Hubinger told his followers that “Jacob is correct here, we really do earnestly believe AI could kill all humans,” and put his personal estimate above 10 percent within the decade. Samuel Marks, who runs scalable oversight at the company, added that developers believe their technology could cause human extinction and that it could happen within a few years. The Washington Post framed this as a crack in the company line. It is closer to the opposite. Anthropic’s market position has always rested on being the lab that says the frightening thing out loud and keeps building anyway, which is why its safety posture has previously made it a target inside the Pentagon rather than a liability with investors.

The Warnings Escalate While the Shipping Accelerates

Set the statements against the release calendar and the pattern is not a debate. It is a ratchet.

WhenThe WarningWhat Happened Anyway
May 2023Hundreds of researchers sign a statement putting AI extinction alongside pandemics and nuclear warCommercial frontier releases continue uninterrupted
February 2026A wave of safety staff exit OpenAI, Anthropic and xAINo lab alters its roadmap
July 2026Capability gains arrive faster than evaluation cyclesAnthropic ships its fourth frontier model in two months
September 2026The alignment lead says there is no plan and no clear track to oneNo pause announced, no roadmap change
Three years of escalating public warnings, set against the release calendar.

Every row contains a warning that was heard, understood, widely covered, and absorbed without altering anything. That is not a communications failure. A warning that changes no behavior over three consecutive years is telling you something about the incentive structure, not about the sincerity of the people issuing it.

The Probability Is the Comfortable Conversation

Here is our read. A p(doom) number is the safest thing an AI executive can say in public. It is unfalsifiable on any timescale that matters, it is a decade out, it flatters the product by describing it as world-ending, and it has never once cost a lab a funding round or a shipping date. Elon Musk spent Thursday calling the whole episode a setup, and the striking thing is how little the accusation matters either way, because a sincere warning and a marketing exercise produce the identical outcome: nothing changes. Hubinger’s admission is different in kind. “We do not have a plan” is checkable, it is about the present, and it is an admission against interest from the person whose job is having the plan.

CNN, September 10, 2026: the segment stays on the extinction question, which is exactly the framing the labs have never had to argue with.

What Is Already Measurable

While the argument runs on 2030, the checkable harms are running now. Stanford economist Erik Brynjolfsson and colleagues, working from ADP payroll records covering millions of workers, found that early-career employment has fallen sharply in the occupations most exposed to generative AI, a relative decline of roughly 16 percent for workers aged 22 to 25 in fields like software engineering, marketing and customer support. The mechanism is not mass layoffs. It is hiring that quietly stops, which removes the on-ramp into those careers without producing a single headline event. No executive has attached a probability to that, because it has already happened and probabilities are for things that have not.

The Question That Has an Answer

Ask whether AI will end humanity and you get a number nobody can score. Ask Anthropic to publish the plan its own alignment lead says does not exist, with dates and thresholds and a named capability level at which the company stops, and the question becomes one that can be checked in twelve months. Axios noted that the extinction debate has broken containment this week. It should break somewhere more useful than a probability argument. The people who confirmed Coxon are not disgruntled leavers with four months of tenure and a grudge. They still work there, they run the safety functions, and they have now said on the record that the safety function does not have a plan. Anthropic should be asked, repeatedly and by name, what it intends to do about that before it ships the next model.