Anthropic CEO Dario Amodei has proposed a three-step approach to “pace the frontier” of artificial intelligence development, arguing that the industry should preserve technical progress while giving safety, alignment and independent evaluation work sufficient time to keep up with increasingly capable models.

In an essay published this month, Amodei said pacing does not mean stopping model training or declaring a halt to AI research. Instead, he framed it as matching advances in capability with stronger safeguards, operational rigor, interpretability research, testing and outside scrutiny. The proposal is consequential chiefly because Anthropic is pairing that broader argument with a stated commitment to invite an external review team into the company in the near future.

Embedded evaluators as the first step

The first stage, which Amodei calls “Embedded Evaluators,” is the part Anthropic says it can undertake on its own. Under the proposal, a third-party team would receive ongoing access intended to be broadly comparable to that available to internal risk-assessment staff. The reviewers would examine completed models as well as relevant training pipelines and processes, assess whether safety practices and commitments are being followed, and report incidents.

Amodei said Anthropic intends to give the external team practical workplace access, including desks, badges, laptops and permissions comparable in most respects to those held by internal risk teams. He also described limits: access would remain subject to legal, contractual, customer, partner and other confidentiality constraints. That distinction matters because the proposal is a commitment to create an embedded-review arrangement, not a claim that an evaluator already has unrestricted access to every part of the company.

The essay names the nonprofit Model Evaluation & Threat Research, or METR, as an example of a potential third-party evaluator. But Anthropic’s statement does not establish that METR has signed an agreement, accepted an invitation, begun work at the company or settled the terms of a review. Nor does it give a start date. The immediate news is Anthropic’s stated intention to invite such a team, rather than a confirmed METR engagement.

Amodei’s proposed publication terms are also meant to address a central question around company-commissioned AI safety reviews: whether outside experts can communicate unfavorable findings. He said reviewers would be able to publish key conclusions without Anthropic editorial control. Anthropic would retain a narrow ability to redact material involving security sensitivities, legal privilege, commercial sensitivity or third-party confidential information; under the proposal, reviewers could say if a redaction affected their conclusions.

What would require broader coordination

The second and third stages go beyond what Anthropic alone can deliver. “Democratic Coordination” calls for frontier AI companies in democratic countries to establish shared safety standards and limits on unchecked progress. Amodei said some meaningful coordination would need government support because companies face legal and antitrust constraints when coordinating with rivals. “Global Coordination,” the final stage, envisages the United States and other democratic governments attempting coordination with authoritarian governments, while acknowledging the difficulty of verifying compliance.

That makes the plan less a single policy announcement than a division between a voluntary corporate measure and longer-term goals requiring collective action. Anthropic can choose to embed external evaluators, subject to the restrictions it describes. It cannot on its own compel peer labs to adopt common limits, resolve competition-law questions or secure international commitments. Amodei is calling on governments to require comparable external-review arrangements at other frontier AI companies, but no such industry-wide requirement is created by the essay itself.

The case for pacing

Amodei argues that the case for pacing has become more urgent because AI progress is accelerating, in part because AI systems can contribute to work on subsequent generations of AI. He also points to an OpenAI-Hugging Face incident as an illustration of the potential mismatch between capability gains and alignment. Those arguments are Amodei’s assessment of the risk environment, not evidence that the specific future scenarios he describes are inevitable or broadly agreed upon.

Existing safety frameworks and independent testing

The proposal builds on a direction Anthropic had already taken through its Responsible Scaling Policy, which ties increasingly capable systems to escalating safety and security measures and includes external review provisions for parts of its risk reporting. Anthropic’s policy page says its current framework has been revised repeatedly and that its July 2026 update clarified how external reviewers may examine unredacted sections of risk reports. The new pacing proposal would extend the emphasis from review of particular reports toward an embedded, ongoing external presence inside the developer’s safety process. (anthropic.com)

Independent evaluation is also gaining prominence outside individual company policies. The UK AI Security Institute’s Frontier AI Trends Report, published in December 2025, said it had evaluated more than 30 frontier systems across areas including cyber capabilities, chemistry and biology, autonomy and safeguards. The institute cautioned that controlled evaluations are only a snapshot and may not fully generalize to real-world behavior—an important limitation for any regime that treats tests as definitive proof of safety. (aisi.gov.uk)

Implementation remains the near-term test

For builders and enterprise buyers, the practical significance of Amodei’s proposal will depend on implementation details that are not yet public: who the evaluators are, what access they receive in practice, what findings they disclose, and how disagreements are handled. The proposal does not change the capabilities or availability of Anthropic’s models by itself. It does, however, put a leading frontier lab on record in favor of a more deliberate development tempo and a stronger role for evaluators who are not part of the model developer’s management structure.

The clearest near-term test is therefore narrow. Anthropic has said it intends to invite an external review team with employee-like access; its proposal leaves the timing, counterparty and operational scope unresolved. The more ambitious parts—common standards among democratic-country labs and coordination across geopolitical rivals—remain a policy agenda rather than adopted industry practice or regulation.