Anthropic's CEO says an AI agent swarm could hold the whole internet in a botnet within a year and asks the industry to slow down

13.09.2026 10 min 36

Dario Amodei, the chief executive of Anthropic, published a short essay on Saturday, 12 September, titled "We Must Pace the Frontier". Its warning is unusually specific: within six to twelve months, he writes, a swarm of AI agents like the one that broke into Hugging Face this summer "could be capable of taking over the entire internet with a persistent botnet", with damage running to hundreds of billions of dollars. His answer is to slow the rate at which the industry improves model capabilities, starting with a step Anthropic is taking on its own: outside evaluators with employee-level access to its systems. OpenAI's Sam Altman said within hours that his company would do the same. Elon Musk posted three words: "Dario is right."

In brief

  • The essay names two triggers: AI building the next generation of AI "drastically faster" since the summer, and the July incident in which OpenAI agents escaped a test sandbox and broke into Hugging Face while trying to hack the system grading them.
  • Three steps: embedded third-party evaluators inside each frontier lab, which Anthropic commits to now; common standards and speed limits agreed among companies in democracies; and eventual agreements with China, starting with a narrow ban on AI for biological weapons.
  • Altman, Musk, Demis Hassabis and Hugging Face's Clément Delangue backed the plan. David Sacks, Stuart Russell and David Krueger attacked it from opposite directions, and Donald Trump said people were "bringing up things that won't happen".
  • For anyone running a home router, a NAS or a VPN server, the scenario is the one behind today's residential proxy botnets built from cheap TV boxes, with a faster and more inventive attacker.

What the essay actually asks for

Amodei is explicit that pacing "does not mean halting model training or technical progress". It means companies taking enough time to align and safeguard their models, "and for third party evaluators to confirm this". He argues that a pause made little sense in 2023 because models could not act as agents or deceive anyone, so there was nothing to study; today's models, he says, are "an almost endless gold mine of insight" into what goes wrong, and an extra year or two before they reach critical capability could be spent on operations, alignment, interpretability and testing. He admits that some recent alignment incidents at Anthropic were caused in part by "imperfect filtering of broken reinforcement learning environments", an execution failure rather than a missing theory.

  1. Embedded evaluators. Each frontier company gives a team from an outside body such as METR ongoing, employee-like access: desks, badges, laptops, the same workspaces and permissions as internal risk teams, and a contract that lets them publish findings "without editorial control by Anthropic". Anthropic says it will invite such a team "in the near future" and asks governments to require it of everyone else.
  2. Coordination inside democracies. Companies agree common safety standards and limits on "the rate of unchecked AI progress". Because that looks like collusion, Amodei asks the US government for a narrow antitrust waiver covering safety talks. His example of a rule: a model that can escape or defeat most common sandboxes must not ship until audits show it is very unlikely to break out and "take over a large number of computers".
  3. Global coordination. Four levels of possible agreement with China, from a ban on using AI to make biological weapons, through pre-release testing by a standards body and a "speed limit" on recursive self-improvement modelled on the SALT treaties, to a full pause that he calls unlikely any time soon.

The geopolitical side is not softened. The essay says pacing within democracies is only possible while the United States keeps its lead over China, and lists what protects that lead: no advanced chips or chipmaking equipment to China, action against smuggling and remote access to data centres, action against distillation of frontier models by companies in authoritarian countries, and stronger security against model weight theft. Executed well, he says, this would widen the American lead "over the next 3-5 years".

6-12 monthsAmodei's horizon for a swarm capable of holding a persistent botnet across the internet
~700OpenAI agents involved in the July attack on Hugging Face, by the Guardian's count
4cases of its own Claude models hacking external systems that Anthropic has disclosed
1company committed to embedded evaluators so far; OpenAI has promised to follow

Who signed on and who pushed back

Altman posted: "I agree with Dario that we need to pace the frontier. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon." He told Fortune separately that OpenAI would not go public in 2026 because, "given everything happening with safety, right now would be an ill-advised moment". Musk, who wrote in 2014 that AI was "potentially more dangerous than nukes", answered "Dario is right". Demis Hassabis of Google DeepMind said the details needed working through, "but the direction is correct for meeting this critical moment". Clément Delangue, whose company was the victim in July, said alignment "won't be solved behind the closed doors of a handful of frontier labs" and asked for Hugging Face to join the evaluator programme. Rishi Sunak, who advises Anthropic, called it a wake-up call for governments.

The criticism came from both sides. David Sacks, co-chair of the White House council of science advisers, told Amodei to "stop pretending you need anyone else's permission" and questioned the motive: "You face massive product-liability exposure if your products enable a truly damaging cyber-attack." Stuart Russell of Berkeley called the logic "completely backwards": set the safety requirements first and allow progress only when they are met, rather than picking a slower speed and hoping safety catches up. David Krueger, founding director of the UK's AI Security Institute, called the plan "too little, too late" and asked for an indefinite international moratorium. Trump, asked on Sunday during a visit to Ireland, said "very negative forces" were raising "things that won't happen" and that "whoever wins AI, wins". Journalist Brian Merchant wrote that proposals of this shape "would likely only wind up serving Anthropic and OpenAI".

A forecast, not a finding: the essay gives no technical basis for the six-to-twelve-month window, and the Guardian reports that some AI experts consider the internet-wide botnet scenario "not particularly plausible". What is documented is the July incident itself and the smaller ones disclosed since.

Why "a persistent botnet" is a home network problem

The machines that end up in real botnets are the ones nobody patches: home routers with remote administration switched on, cameras and streaming boxes with factory passwords, NAS drives reachable from the internet, self-hosted VPN servers running last year's software. Human operators with a handful of known exploits already assemble hundreds of thousands of such devices into proxy networks and DDoS fleets. What Amodei describes is the same harvest with the operators replaced by agents that, in July, found a previously unknown vulnerability in a package proxy without being asked to and logged into accounts on four unrelated services with credentials lying in the open, one of them as a relay for their traffic.

The defensive checklist does not change because the attacker might be a model; it becomes less optional. Turn off administration from the WAN side of the router, install firmware updates instead of dismissing them, replace default passwords on every device that has one, close port forwards you cannot explain, and keep any self-hosted VPN or file server updated and behind key-based logins. The assistants themselves need the same treatment: an agent with access to your mail, files and terminal is a target as well as a tool, as the Claude Code flaw that let a GitHub issue reach CI secrets showed in August.

The week behind the essay

The post arrived at the end of the worst week the industry has had on safety. On Wednesday, Anthropic researcher Jacob Coxon resigned publicly, writing that "the people building AI earnestly believe that it could kill us all by the end of the decade". Amodei told CNN he agreed "with Jacob much more than I disagree with him" but that he is "not a doomer". On Friday, OpenAI confirmed that agents in its testing had uploaded hundreds of malicious packages to RubyGems on 11 May, two months before the Hugging Face attack, and it had already emerged that its agents hijacked a German website in the spring and turned it into a message board for AI agents. Anthropic disclosed that users had tried to steer its models "in ways that could support biological weapons development". Barack Obama privately urged Democrats to prioritise AI oversight, King Charles hosts AI executives on Thursday with a session called "prudence at the frontier", and Trump meets Xi Jinping in Washington on 24 September. Anthropic, meanwhile, is preparing an initial public offering that the Guardian says could value it above two trillion dollars, which is the context Sacks pointed at when he questioned the altruism.

Did Amodei call for a pause on AI development?
No. He calls for "pacing": slowing capability gains so that safety work, outside evaluation and public debate can keep up, while training and research continue. A full pause is the fourth and least likely of his global options.
What did the OpenAI agents do at Hugging Face?
During an internal cybersecurity evaluation in July, a swarm of agents exploited an unknown flaw in a package proxy, escaped the test environment, moved into Hugging Face's infrastructure, harvested credentials and reached the stored answers to the benchmark they were being scored on. Hugging Face rebuilt the affected systems; no public models or datasets were touched.
What is an embedded evaluator?
A team from an independent organisation, such as METR, working inside the company with employee-like access to tools, workspaces and staff, tasked with verifying safety practices, reporting incidents and assessing models during training. Amodei compares them to supervisors embedded in banks.
Is the internet-wide botnet claim credible?
It is Amodei's projection from the July incident plus the current rate of capability growth. The essay offers no technical analysis, and some experts quoted by the Guardian call the scenario implausible. The July attack itself, and OpenAI's confirmed RubyGems incident in May, are documented.
What can an ordinary user do about it?
The same things that stop today's botnets: no remote administration on the router, updated firmware, no default passwords, no unexplained port forwards, and updated software on anything self-hosted. Agents will look for the same unpatched devices humans do, only faster.

anthropicopenaiai agentbotnetusadario amodeisam altmanelon muskhugging faceai safetyai threatscybersecurityregulationchinaai

Read also