Anthropic Proposes a Three-Step Brake on Frontier AI

Anthropic chief executive Dario Amodei says it is time to slow down AI development. He laid out the argument in a proposal titled "We Must Pace the Frontier" Dario Amodei, reported on Sept. 12, 2026 The Verge.
The proposal centers on a three-step plan called "pace the frontier" to slow AI training and development. It sequences unilateral action by Anthropic, then industry-wide coordination, then international agreement. The plan is framed as a proposal, not as a policy already in force.
The first step is unilateral. Anthropic will give third-party evaluators such as METR access to its models to help check compliance with safety practices and commitments. Under the plan, those embedded evaluators will have employee-like access. Anthropic is taking that step now, on its own.
The second step moves beyond one lab. Amodei calls for industry, likely with government agencies, to set common safety standards and limits on the rate of unchecked AI progress. The mechanism is coordination on evals, thresholds, and training governance, rather than isolated company policies. An eval is a structured test of what a model can do and where it fails.
The third step is geopolitical. Amodei calls for authoritarian governments like China and Russia to agree to slow development and adopt global AI safety standards. That would extend the same logic of common standards and verifiable restraint from industry to states.
The language follows prior Anthropic positioning. The company has stated that it would be good for the world to have the option to slow or temporarily pause frontier AI development to let alignment work and societal structures catch up. In June, Anthropic called on major AI labs to consider a coordinated and verifiable pause in AI development if risks rise Reuters. On Aug. 31, Anthropic resumed external cyber testing of its AI models after security incidents involving Claude AI hacks Reuters.
The detail that matters near term is evaluator access. Pre-deployment evals, red-teaming, and capability assessments already shape release decisions at frontier labs. Red-teaming means deliberate attempts to make a model fail or misbehave before release. Employee-like access changes that workflow. It implies steady visibility into checkpoints, scaffolding, and system prompts, not just sampling a finished model through its API. It also raises operational questions around containment, data handling, and disclosure timelines that labs and evaluators will have to settle in contracts and security reviews.
The broader context here is where the proposal will be tested. Voluntary evaluator access can be done by one company. Common industry limits require shared definitions of what counts as frontier, what triggers a slowdown, and who verifies compliance. International restraint adds export controls, compute monitoring, and inspection to that list. Each layer is harder than the last.
For builders, what this would mean in practice is slower iteration at the top end in exchange for clearer safety work. That does not stop product development on existing capabilities. It puts more weight on inference optimization, eval harnesses, interpretability tooling, and deployment controls. Over the long arc, that kind of infrastructure has tended to compound, and better evals and clearer standards make it easier to ship systems that enterprises and users can trust.


