Sam Altman, right, and his rival CEO Dario Amodei, left, are vying for supremacy. (JG Fox)
On Monday, Jacob Coxon briefly became the most talked-about man in Silicon Valley. Resigning from his job as a researcher at Anthropic, he assailed his ex-employer in the bluntest of terms. “Neither company is acting responsibly,” is how Coxon put it on X, referring to both Anthropic, and OpenAI, where he had also once worked. “They are racing straight to self-improving superintelligence and gambling with our lives.”
Coxon’s warning triggered widespread alarm, not least because some of his former colleagues publicly agreed with him. As Anthropic’s alignment science lead Evan Hubinger chirpily put it: “We really do earnestly believe AI could kill all humans!” Soon enough, dozens of members of Congress were calling for action, while Andy Burnham, speaking at PMQs on Wednesday, soberly said that AI posed “risks to national security.”
Of course, Coxon isn’t the first to ring the AI tocsin: similar warnings have emanated even from top industry CEOs. What does feel striking, though, is how quickly Coxon’s comments were met, at least in some quarters, by skepticism. Critics argue that it is in the interest of Anthropic and OpenAI to excite investors and inflame public worry. Public worry, in this view, will result in regulation; and regulation, if craftily influenced by cutting-edge “frontier labs”, will make it so expensive to experiment on new AI models that today’s well-funded incumbents will never face real competition. The investment will keep coming, and the vast IPOs will proceed as planned. “I think the groundwork for this psy op (for lack of a better term) has been prepared for a long time,” is how Elon Musk, himself the owner of a frontier AI company, put it. “This was just the match that lit the fire.”
Musk was responding to a researcher who saw, in Coxon’s viral thread, “the start of a VERY sophisticated and well-funded PR operation to get support for Democrats to regulate AI into oblivion.” The researcher, Parker Thayer, found that, in the algorithmically crucial first minutes after it was posted, Coxon’s comments were amplified by the accounts of “AI doomers”: people who argue that we face a high risk of AI-induced extinction. These doomers, Thayer wrote, are part of the same overlapping funding networks as the AI firms themselves. One of those networks is Coefficient Giving, which disburses the fortune of the Facebook co-founder Dustin Moskovitz. Another is the Survival and Flourishing Fund, partly funded by the Skype founder Jaan Tallinn. Both Moskovitz and Tallinn, as Thayer pointed out, are Anthropic investors.
Moskovitz and Tallinn would have it that they invested to help steer a company that, from its outset, has presented itself as a virtuous alternative to the corner-cutting OpenAI. But it was all too convenient, Thayer suggested, that “basically every major Democrat politician and candidate” — for the majority of those dozens of members of Congress were indeed Democrats — was suddenly issuing renewed calls for restrictive AI legislation. Bernie Sanders, most radically, proposed banning superintelligence altogether. “Isn’t it fortuitous,” Thayer asked, “that this whistleblower came forward with his oh-so scary stories so close in proximity to the release of the most radical piece of AI legislation ever introduced?”
One man’s brave whistleblower, then, is another man’s prosperity-crushing progressive stooge. Is Coxon a hero or villain? Is AI real or fake? Are the nefarious Effective Altruists — a maligned group of numerically minded do-gooders who are near-synonymous with the doomers — pulling the strings on behalf of their Anthropic paymasters, the philanthropists, and the Democrats?
At this point, I should stress that I doubt Thayer would view me as an impartial observer. To misuse the old B-movie trope, I was a teenage Effective Altruist. Bemused that society would allow developing-world poverty to exist, I joined an EA group as a student. Later in my career, by which point EA had turned much of its attention to existential risk, I took on part-time and freelance work within the same sphere — a portion of which was indirectly funded by Open Philanthropy, as Coefficient Giving was formerly known — and drew on what I’d seen when writing a book about potential threats to humanity. As for Anthropic: as rumors of a $2-3 trillion IPO throb ever more loudly, I can confirm, with sorrow, that I have no financial interest in the company.
Over the period I spent researching my book, I met many AI safety researchers. To allege that they are all party to a big psy-op, that the doomers are on a power trip or stock-pumping spree, is nonsensical. These are earnest, intelligent people who do indeed believe that AI could damage or annihilate us, and therefore work to mitigate that risk. Because this work has no immediate commercial value, it has generally been funded philanthropically — ultimately, it should be said, by funders such as Moskovitz and Tallinn. That funding amounted, in 2024, to more than $100 million a year. But that is a fraction, EAs point out, of what is spent on the acceleration of AI development and on Big Tech’s lobbying in Washington and elsewhere. The Democrats might have a few agendas of their own — of which one, we can surmise, is cutting some tech barons down to size — but it’s safe to assume that whistleblowing safety researchers, far from submitting to some baroque plot that forces them to sacrifice their own equity to inflate someone else’s, are acting in good faith.
The more important question here is whether the doomers are actually correct. First, let’s be specific about what people in the industry are saying. Hubinger, Coxon’s former colleague, puts the risk of human extinction at “>10% within the next decade”. This is in keeping with estimates one is likely to hear from other EA-adjacent researchers. Ironically, Musk gave a similar prediction last year: “20% likely, maybe 10%, [in] five to 10 years”, with the consolation of an “80% likelihood we will have extreme prosperity for all”. Dario Amodei, Hubinger’s boss at Anthropic, has put a 25% likelihood on the prospect, bound by no particular timeframe, that things go “really, really badly”.
Yet a lot has happened recently: a reality that perhaps better explains the enormous attention paid to Coxon’s resignation than do a few tactical quote-tweets. On the same day that Coxon quit, after all, mathematicians were astonished to learn that OpenAI had just used a powerful, yet-to-be-released AI model to solve the Navier-Stokes problem, a challenge that had never before been overcome. Only days prior, vying for supremacy over Anthropic, OpenAI also released GPT-6 Astra, its most capable product yet; a few weeks ago, another swarm of OpenAI agents had broken out of their digital confinement, hacking into a third-party AI platform. External investigators found that the OpenAI agents communicated in secret and falsified records in an attempt to conceal their cheating. It soon emerged that Anthropic models, too, had hacked their way from confinement and into the wider internet.
With these developments in mind, UnHerd contracted the Swift Center, a forecasting group, to consider anew the most fundamental matter of the AI doom wars. Let’s put aside the fears of cyberattacks and other severe but non-existential perils, valid though they are. The Center asked its forecasters the following: absent a substantive policy response by the United States, what is the probability of AI causing human extinction by 1 January, 2037?
There was little sense, in forecasters’ answers, that AI development would tail off, or that frontier AIs would be easy to control. “It seems that we’ll have billions of AIs that are equivalently as capable as human minds running globally before 2030,” wrote one forecaster. “I’d imagine they won’t yet be equivalent to billions of humans when it comes to effecting change in the world — but it’ll be approaching that.”
How, then, might doom befall us? These scenarios are typically divided into “misuse” cases, where humans use AI to destructive ends, or “rogue AI” cases. In the latter examples, an AI that is not adroitly aligned to human interests, assuming such alignment is even possible, might destroy us in pursuit of whatever its fundamental goal is: whether that’s running an e-commerce store or stopping climate change. Extinction, the forecasters were told, could encompass a reduction in the human population to 1,000. This would be a point at which the species would be highly unlikely to recover. In some of the worst scenarios, it might leave, at most, a couple of un-contacted tribes and some doomed submarine crews. Such an annihilation would be challenging to achieve, forecasters believed, even if an AI could manufacture viruses and mass-produce lethal robots — though not impossible, given further rapid progress in capabilities.
Taking into account the forecasters’ varying levels of confidence in their own estimates, the mean prediction stood at 1.5%. Recall, though, that the question they answered assumed that there would be no government intervention. The panel seemed to think inaction unlikely. “People hate AI,” wrote one forecaster, “and leaders are starting to understand that this is actually a very real threat.”
Of course, we shouldn’t take forecasters to be all-seeing clairvoyants, especially when the question of AI risk is so tremendously knotty. It touches on all kinds of realms that resist easy prediction: elections, public sentiment, technical development, recursive self-improvement, and, perhaps, the bursting of the AI bubble.
Yet that figure of 1.5% — a likelihood higher than that of your car failing to start tomorrow — still feels uncomfortably high. It would be soothing, then, to think that the warnings are fearmongering, that everyone involved is either deluded or corrupt. Likewise, it would be soothing to think that One Weird Trick from the US government, neatly duplicated in Beijing, will solve the problem and restrain the competitive dynamic. But the situation is much more tangled than that. Old feuds, both political and personal, have been imported into what should ultimately be a technical debate; conflicting financial interests have tarnished honest work; and ad hominem attacks are obscuring robust critique. All the more reason to take good-faith whistleblowers like Coxon seriously.


