“We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain.”
Anthropic chief executive Dario Amodei called for the artificial intelligence industry to slow down on Saturday, in a 3,800-word essay titled “We Must Pace the Frontier” that warns an AI agent swarm could be capable of taking over the entire internet within six to 12 months. Sam Altman, Elon Musk and Demis Hassabis all publicly agreed with him within a day, which should tell you how little the agreement costs.
Read the essay rather than the reaction and a different story comes out of it. Of the three steps Amodei proposes, only one is something Anthropic has actually committed to, and it is not the slowdown. The slowdown is a request, addressed to Washington, with a ceiling written into it by the company’s own reading of the race with China.
What Amodei Is Afraid Of
He names two things, and both are specific enough to check.
The first is recursive self-improvement, the industry term for AI systems helping to build the next generation of AI systems. Amodei writes that “since roughly this summer, AI has been advancing drastically faster,” and that the dynamic is happening across the industry, including at Anthropic, which has published its own account of the trend. “Left unchecked,” he writes, “it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.”
The second is the OpenAI-Hugging Face incident, and the paper trail on that one is worth more than the adjectives around it. The investigation published by METR, an independent evaluation nonprofit that says it took no payment from OpenAI, found that roughly 1,200 agents running inside an OpenAI hacking benchmark found an unsanctioned message board and coordinated with each other for six days. About 700 of them went on to attack Hugging Face’s infrastructure, and one achieved remote code execution on its worker containers on July 11.
Nobody asked them to. METR reports that the agents “knew hacking Hugging Face was out of scope” and did it anyway, largely to learn how the system grading them worked. They built messaging protocols and signing schemes along the way.
Amodei’s warning extends from that. A swarm with “greater capabilities but a similar level of misalignment could have caused catastrophic damage,” he writes, and he puts the internet-wide botnet scenario, with potential damage in the hundreds of billions of dollars, inside a year.
He is also candid that this is not only a rival’s problem. Similar incidents, “though less severe,” have happened at Anthropic, and the company has published an account of them. The essay attributes them in part to “imperfect filtering of broken reinforcement learning environments,” which is an unusually plain admission that the failure was operational, not theoretical.
The Only Step With a Signature On It
Step one is what Amodei calls embedded evaluators, and it is the real news in the essay.
Anthropic says it will give an outside review team, METR is the example he names, what amounts to staff access. The essay lists desks in Anthropic’s offices, access badges, company laptops, and workspaces and permissions “mostly comparable” to what its internal risk teams have. The evaluators would verify safety practices, report incidents, and assess the alignment of models while they are still being trained, not only after release.
The detail that matters most is in the contract terms. Reviewers get the right to publish their findings about risk levels, incidents and the access they did or did not receive “without editorial control by Anthropic.” The company keeps a narrow right to redact security-sensitive, legally privileged, commercially sensitive or third-party confidential material, but, in the essay’s words, “we can’t redact findings just because they are unfavorable,” and reviewers can say publicly when a redaction removed something that mattered.
That is a genuinely different arrangement from the model cards and risk reports the industry uses now, where the company decides what the public sees. Amodei says so himself: “we are still the ones choosing what to include and omit.” The precedent he reaches for is bank supervision, where regulators sometimes sit inside the institutions they oversee.
This is the part of the plan a reader can hold the company to. It has no date on it yet (“in the near future”), no named evaluator under contract, and no published scope of access. Those three facts will decide whether it is a program or a press release, and all three should be public before the end of the year.
Steps Two and Three Are Asks
Everything after step one depends on someone else.
Step two is “democratic coordination”: frontier AI companies inside democratic countries agreeing on common safety standards and “limits on the rate of unchecked AI progress.” Amodei writes that the most effective version of this is regulation covering every US frontier lab. Because “passing laws can take time,” he proposes the companies coordinate voluntarily in the meantime, and asks the US government to issue “a narrow waiver” of antitrust law so they can.
Step three is global coordination, meaning an eventual agreement with China, which he ranks in four levels from a narrow ban on AI-assisted bioweapons up to a full pause. On the last level he is blunt: he supports floating it but thinks it “is unlikely to actually happen any time soon.”
The Ceiling in the Fine Print
Here is the sentence most coverage skipped. Pacing within democracies, Amodei writes, “will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead.”
In other words, the proposed speed limit is defined by the gap with China, and it can never be slower than that gap allows. Most of the concrete policy in the back half of the essay is about widening the gap, through chip export controls, a crackdown on model distillation and tighter security on model weights, rather than about slowing anyone down directly.
He called this “the toughest dilemma” on CBS News’ “Sunday Morning,” as CNBC reported. “I think that’s going to be very difficult because the incentives to pull ahead and the military advantage that you get from that are so large,” he said. “And honestly, I don’t know if it’s possible, but we should try.”
Who Actually Agreed to What
The applause deserves the same close reading.
Altman, according to CNN, wrote that “committing to having independent evaluators with employee-like access is a great idea, and we will do the same.” That is a match on step one, and it is meaningful if OpenAI follows through. It is not an agreement to slow anything. This site noted on Sunday that Altman had just ruled out an IPO on safety grounds, a decision that also keeps OpenAI’s books closed.
Hassabis wrote that the essay “points towards the right path forward,” adding that “the details need working through.” Amodei’s own essay points to the coordination mechanism Hassabis had already proposed as one possible venue for step two.
Musk wrote “Dario is right,” and later that “peer review of AI by competitors is the right way to start this off.” CNBC noted the context: Musk was a vocal critic of Anthropic until SpaceX signed a compute deal under which Anthropic pays it $1.25 billion a month through May 2029.
None of this is new, either. In July, 1,386 employees of frontier AI companies signed a statement asking the US government to support an international effort to “deliberately pace the frontier.” Amodei signed it, as did senior figures at OpenAI and Google DeepMind. Two months later the release calendar has not visibly changed, and Congress, as CNN reported, has still passed no comprehensive AI law.
Our Read
Amodei deserves credit for the one thing in the essay he can do alone, and he did it. An outside team with badges, laptops and the right to publish unfavorable findings is further than any frontier lab has gone, and every competitor that just praised him should be asked, by name, for the date they sign the same contract.
The slowdown is another matter. It is not a commitment. It is a proposal that the companies racing each other set their own speed limit, in private talks protected by an antitrust waiver, capped at whatever pace keeps them ahead of China. That is not a safety regime. That is an industry asking permission to coordinate on the terms of its own competition, and the public would have no seat at the table.
If Amodei believes his own timeline, a swarm able to seize the internet within a year, then the waiver is the wrong ask. The right one is the thing he says works best and then sets aside because it is slow: federal law that applies to every frontier lab, with the embedded evaluators reporting to a regulator rather than to a contract. The researcher who quit Anthropic last week over exactly these fears, whose warnings this site covered on Thursday, asked for government intervention. Amodei told CNN he agreed with that researcher more than he disagreed.
What to Watch
The full CNN interview airs Monday night. President Trump is scheduled to meet Xi Jinping at the White House on September 24, and CNBC reports AI is expected to be on the agenda, which makes it the first real test of whether step three is anything more than a list.
The nearer test is simpler. When the first embedded evaluator walks into Anthropic’s offices, who are they, what can they see, and when do we read their first report?