Anthropic promises embedded evaluators as Amodei urges slower AI gains
Anthropic will give third-party evaluators ongoing, employee-like access to its training and safety systems, chief executive Dario Amodei wrote Saturday. The team would verify safety practices, report incidents and assess models as well as training processes.
The commitment is immediate. Amodei's next two steps would set common limits among frontier labs in democratic countries, then seek verifiable agreements with authoritarian governments. Both would require coordination beyond Anthropic.
The prediction is his. Amodei said a more capable misaligned agent swarm could within six to 12 months sustain an internet-wide botnet and potentially cause hundreds of billions of dollars in damage. He called for slower capability gains, not a halt to model training.
The document: Dario Amodei, We Must Pace the Frontier, September 2026.

