September 14 2026 - 10:00am

The chief executives of America’s leading AI developers rarely show public support for one another. Elon Musk co-founded OpenAI but is now a bitter critic of “Scam Altman”, the company’s CEO. Dario Amodei’s Anthropic was founded by former OpenAI staff who argued that Altman was commercializing AI too fast. Musk’s xAI, now merged with SpaceX, was founded to rival them both.

But this week has been no ordinary one in AI. The resignation of Anthropic researcher Jacob Coxon, whose warnings of an AI-induced extinction event were echoed by former colleagues, sparked widespread concern — and the loudest calls yet for the US government to step in. Altman, Amodei and the rest are racing to the frontier, argued the fearful, and in their haste are endangering us all.

Amodei responded on Saturday by making an intriguing public commitment: to unilaterally embed third-party safety evaluators within his own organization. “Evals” teams already work with frontier labs, but on a limited, precarious and often remote basis. If frontier AI labs are trying, as one anonymous OpenAI employee put it, to safely breed a digital Tyrannosaurus rex, evals teams check that the cage is robust and see what the dinosaur does when given the chance to misbehave.

Anthropic, said Amodei, would invite an evals team into its offices, giving it desks, access badges and access to the company’s training and safety procedures. Evaluators, he pledged, would have the right to publish, without editorial control from Anthropic, “key findings about risk levels, incidents, practices, and the access they received or didn’t receive”.

Within hours, Musk and Altman expressed support for the plan. Both CEOs said they were willing to meet Amodei and other frontier-lab representatives.

In this fractious and competitive industry, safety initiatives have not always been seen through. In 2023, for instance, OpenAI announced the creation of a “Superalignment” unit. The unit was dissolved within a year.

One factor that could make Amodei’s proposal more durable is that it might keep legislators off the labs’ backs. Mindful, presumably, of the growing appetite in Washington, DC, for action, he argued that self-regulation would be more appropriate than slow-moving government intervention.

In doing so, he has drawn the ire of David Sacks, who is a co-chair of President Trump’s Council of Advisors on Science & Technology. “Stop pretending you need anyone else’s permission,” he wrote, casting doubt on Amodei’s suggestion that AI might soon bring down the internet. “Stop pretending antitrust law has to be suspended so you can form a cartel.”

If the plan is indeed durable, it will vest considerable power in whichever evals organizations are invited into the frontier labs. Some of them are relatively small philanthropically funded nonprofits partly staffed by frontier lab resignees. Others are government departments. Of these, Britain’s AI Security Institute (AISI), far better funded than its American counterpart, is regarded as the most effective.

More evals organizations could spring up. If they work with the major labs, all will have their competence, independence and motives scrutinized by critics.

Since its inception in 2023, AISI has been given early access to frontier labs’ forthcoming models. Recently, though, for reasons that remain publicly unclear, it was barred from conducting an evaluation of the Anthropic model Mythos.

But Amodei has also proposed that the US seek to cooperate with China, with his proposals ranging from an agreement not to use AI to create biological weapons (relatively straightforward) to a pause in the development of powerful AI. Evals, then, could have a diplomatic dimension, much as nuclear weapons treaties typically require their armed signatories to submit to regular inspections by neutral third parties.

Already, there has been discussion in government circles of the role that AISI could play in a world of embedded evals. The ability to help shape or slow AI is not remotely as significant as the role of developing it, but it is better than no role at all.

Should Amodei and his counterparts be trusted? One skeptical reading of his proposals is that they amount to a pretext for Anthropic failing to live up to its promises. Another, the one articulated by Sacks, is that the scheme will sabotage competitors by making it onerously expensive or bureaucratic to develop frontier AI.

The anonymous OpenAI employee, the one who likened frontier model training to breeding a T-Rex, had a different view. Once the labs have learned, at great expense, to breed the creature safely, he argued, “it’s far simpler for a new entrant to create, contain and secure a Stegosaurus.”

The viability of any lasting accord among Western labs, however, will depend on the relationship between the US and China. Neither country wants to lose the AI race, a rivalry that will inevitably affect relations between the American labs. A problem in this respect is that the capability of Chinese models relative to their American counterparts remains unclear. Does the US believe Xi Jinping is breeding T-Rexes? Or just stegosauruses?


Tom Ough is a senior editor at UnHerd. He is the author of a book on existential risk.