Ternus, the new Apple CEO, speaking on Wednesday. (Benjamin Fanjoy/Getty)
The physical mobile phone seems to have stagnated. Take the next iPhone, which new Apple CEO John Ternus has just unveiled. Featuring a folding screen and incrementally improved processing, that might excite Apple’s legions of fans, even if they have to pay $2,000 for the privilege. All the same, Ternus’ announcements ultimately constitute a facelift rather than a revolution, with the upgrades a far cry from the boom of innovation we saw in the years after the iPhone was first launched. Instead, like its great rival Google, Apple is increasingly building its products around agentic AI, the kind of software that can be a personal assistant, a playmate, a business partner.
Yet the real paradigm change here isn’t AI-as-agent, but rather the shift to a phone as the main mechanism for the use and development of AI itself. Your next phone, be it a new iPhone or a new Google Pixel, be it iOS or Android, will be a figurative microphone onto the world — a microphone from which AI models can collect ever-increasing amounts of data to learn from, shaped in private by methods the public only has the foggiest idea of.
What the public does surely understand is that vast resources are being poured into the building of AI. This is done partly through the building of AI infrastructure itself: creating LLM architectures that can contain trillions of parameters and run on hundreds of thousands of graphics-processing units (GPUs). That in itself is no small investment of time and energy, as evinced by the decades of promising but unprofitable research required before the explosion of Deep Learning in the 2010s.
What’s sometimes overlooked, though, is the actual training of the models: the process by which these architectures are fed enormous amounts of data in order that they are not only useful, but are useful in the “right” way. And here is where the AI titans — OpenAI, Anthropic, Google — have a seemingly permanent advantage that no open-source model or start-up could hope to match.
AI companies tend to be quiet about where they source their data from. After the great fuss made in the 2000s over the downloading of single music files, and the endless attempt to shut down SciHub, the repository of pirated scientific research, it was ironic to learn that Facebook had downloaded millions of books from pirate site Library Genesis to train its AI. The kicker was when a federal judge found, when judging in favor of Meta, that such piracy could well qualify as “fair use”, since the results of it were “transformative”.
The point here is less about what big companies can get away with, and more about the centrality of training data to the capabilities of AIs. We’re beginning to see a push for people to upload their health data to LLMs, with OpenAI offering its American users a feature that analyzes their health records. There are all sorts of questions raised around this: could private data surface in public, even after being anonymized? Where will liability for wrong answers fall? How prescriptive should such an agent be when it comes to lifestyle and health? And should a handful of companies be assembling the models that wield this sort of influence without much in the way of oversight?
Some dangers are already clear. There are the recurrent hints that these companies don’t have these powerful and opaque models under sufficient control. In the recent past, both OpenAI and Anthropic’s models hacked out of their own sandboxes and into each other’s secure infrastructure, the sort of thing that you’d be forgiven for expecting AI companies to know how to prevent by this point. Just yesterday, the Anthropic researcher Jacob Coxon resigned, saying his employer and OpenAI were “gambling with our lives”, only for his ex-colleague Evan Hubinger to chime in and agree: “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” Touché?
Those who fear that advanced AI will bring catastrophic consequences can point to the fact that it is still quite commonplace to “jailbreak” even the most advanced AIs to get them to cough up data and responses that they’re not supposed to. For example, on the day GPT-6 (Astra) was released, the researcher Sergey Berezin was able to get it to name pirate sites like ThePirateBay with no copyright disclaimer.
The jailbreaking issue, however, touches on a less-treated subject. No AI conclusively learns “on its own” to refrain from talking about pirate websites, or responding with pornographic content. Nor, for that matter, does it understand from the womb that it should not hack out of its sandbox. All such guardrails are inserted after the fact, most often through a murky process of “post-tuning” and the addition of further guardrails. In both cases, AI companies have an undisclosed process of human review. As OpenAI describes it, they instruct their reviewers to use their “expertise” to test whether the AI models “do not complete requests for illegal content” and “avoid taking a position on controversial topics.”
Behind this somewhat vague and innocuous-sounding text is a huge amount of influence over the behavior and dispositions of AIs. By this point, it’s accepted that Google, Facebook, Instagram, YouTube and other social content platforms exert a great deal of moderation to hide or remove content that is legal but nevertheless goes against what these companies deem to be community norms. It’s generally agreed that a wholly unmoderated public internet is unworkable. But AI represents an even more profound form of interference.
Instead of moderating content piecemeal, AIs are, by and large, enormous black boxes. Every piece of an AI model, in theory, might affect any other, and so AI companies are engaged in the creation of online intelligences that do a tremendous amount of selection and shaping over the knowledge they synthesize. While the frequent unpredictability of the AIs’ behavior is worrying, so too is the shaping of AI opinion by such a small number of players. OpenAI may gesture at their anonymous reviewers’ unspecified “expertise”, but the fact remains that the process is internal, without oversight, and ultimately controlled by the company.
The results are surely impressive, besting human performance a good deal of the time. OpenAI’s solving of the Navier-Stokes problem, completing one of the great quests of modern mathematics, is just one of many examples. But these successes should not blind us to the possible pitfalls of the way LLMs are trained — especially given that whatever skewing takes place in the post-tuning process is considerably harder to detect than Google removing a particular site from its search engines, or Facebook banning a particular post. These companies are shaping our collective consciousness in a far more wide-ranging and less-noticeable way.
Earlier this year, a reporter found some suggestive evidence that chatbots tend to nudge people away from political extremes. That tendency could stem from an intentional policy on the part of AI companies or, more likely, it could be simply a consequence of post-training AIs to be more anodyne and less controversial. Either way, the fact is that the AIs do increasingly exert such influence. It may indeed be good for societal cohesion for people to have an AI friend that generally pushes them away from dogmatic extremes. But should the construction and maintenance of those friends be operated more or less in the dark by a handful of companies?
What we’re seeing, in fact, is a less intrusive and more benign version of China’s army of censors — the apparatchiks who police online content within what the CCP deems permissible on its “internet”. There’s no such explicit policing in the construction of AI models; but, however benign, it is the concentration of such censorious authority in a handful of non-governmental entities that’s concerning.
That’s only one side of the coin. The flipside is that even as these companies construct these artificial agents of influence, they are simultaneously taking in new data that is exclusively theirs. With every iteration, AIs draw more and more of their training from their predecessors, and the everyday interactions people have with AIs — treating them as friends, business advisers, confidantes, and occasionally intimate partners — provide crucial data for how best to appeal even more to users, win their trust, and be even more convincingly human.
Having already digested the entirety of the internet’s publicly-available data, future AI models will increasingly contain the residue of countless private interactions between users and the AIs’ predecessors. For the time being, it seems that ChatGPT could hold on to all chats forever unless you explicitly opt-out, as does Claude, even if the picture is foggier with Gemini, which is Google’s model. That in itself is a privacy nightmare. But even if the companies were to delete the raw chats, some great portion of your data will remain incomprehensibly echoed in the dark parameters of the next generation of AIs, and in every subsequent generation. They will display an uncanny knowledge of humanity large and small, not traceable to any particular source beyond the totality of their private training data. As AI becomes more proactive rather than passive, the influx of data sucked through your phone to fuel and train AIs will balloon.
Even with sufficient funding and infrastructure, an AI startup today won’t have access to the voluminous chats that existing models have accumulated over recent years. It won’t know what kept ChatGPT’s users engaged, what Anthropic users found useful, or what Claude users found irritating. That feedback is a crucial tool for improving models. Gemini, Claude, and ChatGPT are building an invisible empire of soft knowledge that belongs to them alone, a sophisticated but opaque mirror of the world that gives them an advantage less on technological grounds than on informational ones. When a company like Google, for instance, can control the platform (Android), the network (YouTube), and the intelligence (Gemini), they accumulate a collective package far greater than the sum of its parts and far more difficult to dislodge from prominence.
Apple’s own AI, Siri, is based in pre-LLM technology, and currently lags far behind the competition. The newest version of Siri therefore piggybacks off Gemini — marking Apple as a junior partner. The terms of the agreement between Apple and Google are private, but it is difficult to imagine that Google will allow Apple to make exclusive use of whatever data is gathered. John Ternus, then, has taken the throne of a powerful but stagnant empire, one that is already in hock to technologically superior rivals.
As for AI’s top players, they will end up as de facto self-regulators of the construction and use of AI. Far too useful to ignore, yet too unreliable to wholly trust, AI is now an ever-present problem as much as a boon. And those tasked with managing the problem are currently operating mostly in secret, mostly under the auspices of a few of the world’s largest corporations. The only half-consolation is that even Google and OpenAI and Anthropic do not fully know what it is they are building, and so the immense power they are accumulating is not wholly theirs.



