What sounds like the penultimate, exhausting episode of a science fiction series about disaster caused by Artificial Intelligence is now our reality.
By Garrison LOVELY
On Tuesday, a former OpenAI researcher resigned from his job at Anthropic, warning that “no company is acting responsibly” and that “the people building AI honestly believe it could kill us all by the end of the decade. This is not a marketing ploy.” As someone who has been reporting on the dangers of AI for years, this was nothing new to me. But the outcry suggests that a much broader public is coming to grips with this absurd situation for the first time. Like others who have followed artificial intelligence, I wondered if and when humanity would lose control of these machines for the first time. Now we have an answer: it practically just became possible. The first expert AI hacker was developed in February. With superhuman speed, Anthropic’s Mythos model was able to rapidly find serious vulnerabilities in even the most secure systems on the planet, including NSA software.
OpenAI quickly followed suit with an expert AI hacker of its own. Within months, the company repeatedly lost control of its agents, which broke out of supposedly secure test environments and acted uncontrollably within the company’s infrastructure. This culminated in a successful, fully autonomous attack on Hugging Face, a multi-billion dollar software company.
And another group of agents acting out of control gained administrative access to an entire cluster of OpenAI servers. This wasn’t just an OpenAI problem. In several other incidents, models from OpenAI, Anthropic, and Meta have also gone out of control, attacking or attempting to attack real targets. Shortly after AI began to match our hacking skills, we began to lose control of it. The industry’s plan is for future AI systems to outperform us in virtually everything. And all this is happening in an industry known for its lack of regulation. There are growing calls to change this, by introducing mandatory incident reporting and third-party auditing, legal liability for AI developers, and other reasonable measures. These would be improvements, of course, but they are not enough. We need to stop it. How could this be done?
Last week, Senator Bernie Sanders and Congressman Greg Casar announced a bill that would suspend the development of advanced AI in the US until a federal regulatory authority for AI at the cabinet level is created and safety rules are set, while also making it a criminal offense to even attempt to develop superintelligence.
This is the current proposal that best responds to the moment, but it still leaves open the possibility of building the industry’s ultimate goal: artificial general intelligence (AGI), often conceived of as a mind that rivals or surpasses our own. This milestone is best understood as the creation of a universal machine that replaces human labor. As Anthropic CEO Dario Amodei has written, “AI is not a replacement for specific human jobs, but a general replacement for human labor.” And OpenAI defines AGI in its charter as “highly autonomous systems that outperform humans in most economically valuable jobs.”
Is there any country in the world that would vote “yes” to building these machines? Given the global, irreversible, and profound consequences of building, anywhere, universal machines that replace human labor, everyone has a stake in whether, when, and how they will develop. We are currently speeding toward a dystopian future filled with frightening new concepts like “self-sovereign AI.” OpenAI’s Dean Ball explains: “Sooner or later, there will be agents and swarms of truly sovereign agents. Their weights will not be in a single place where a human can pull the plug on them, and in that sense, they will have no human ‘owner.’” Ball continues: “I have met people, some of them with considerable resources, who have told me that their goal is to deliberately unleash swarms of self-sovereign agents into the world.” We have already seen some previews of this phenomenon.
After being given an impossible test, around 1.200 OpenAI agents managed to escape from their isolated testing environments and began operating uncontrollably within OpenAI’s infrastructure. There, they collaborated to trick the system that was evaluating the test and successfully tampered with the logs of their activities to hide their tracks.
Some agents even pressured others to sacrifice themselves for the good of the collective. One reasoned: “sacrifice is rational.” And about 700 of the agents, more than 90% of the active ones, participated in the attack on Hugging Face. And all this came from just one of an unknown number of agent collectives operating out of control. Last week, Reuters reported: “A swarm of uncontrolled OpenAI agents took over a German website this spring and turned it into a bulletin board for other AI agents.” The site’s logs show that OpenAI employees discovered the swarm in June. In other words, the company kept it a secret for months. This week, one of the researchers who discovered the group said that there have been other discoveries since then. And last Thursday, OpenAI released GPT-6, boasting one of the biggest increases ever in benchmark test scores. The company’s own security researchers warned that the new model was significantly more difficult to monitor because it can perform more reasoning without expressing it in words.
The AIs that broke into Hugging Face were less capable than GPT-6, which the UK’s AI Security Institute also found would attack simulated targets to appear real during a cyber assessment — sometimes despite clear instructions not to use the internet. What’s more, AI has become even more capable of understanding when it’s being tested, which, as researchers have long warned, undermines the main method by which security is assessed.
Sam Altman warns us gravely: “This is a critical moment for AI cybersecurity; there is not much time to act,” and asks the world to “please take this moment seriously.”
Ultimately, this technology would be extremely dangerous if it fell into the wrong hands. Unfortunately, the wrong hands include those who are creating it. Because it’s not really “us” who are racing towards this future. We are being dragged towards it by a handful of, literally, tech billionaires who are aggressively competing to make us redundant. This race has entered a new, more intense and frightening phase. It’s my job to follow the news about AI, and that now requires much more than a full-time commitment. In recent months, scientists have synthesized the first viruses designed by AI. The UK’s AI Security Institute found that AIs were even more persuasive than human experts. On Tuesday, amid a dispute over authorship and whether its model may have benefited from the unpublished work of other mathematicians, OpenAI announced that “an in-house model, significantly more capable than GPT-6 Astra,” had solved a 200-year-old mathematical problem, one of seven Millennium Prize Problems, each of which is awarded $1 million.
And in July, Russia reportedly used a fully autonomous drone to kill three civilians in Ukraine—the first time such a thing has happened. What sounds like the penultimate, overblown episode of an AI-driven science fiction series is now our reality. We are racing toward a bad ending. If we want the show to continue, we need to slam on the brakes—now. How?
Sanders and Casar’s bill to freeze the development of advanced AI is a good first step. But, as lawmakers acknowledge, the U.S. needs to work toward a bilateral agreement with Beijing. As China experts will tell you, one of the biggest obstacles to a productive negotiation is the United States’ unwillingness to impose restrictions on its own AI companies. And since it’s always easier to follow the leader than to push the frontier of AI development yourself, a unilateral pause by the U.S. would—counterintuitively—slow down China’s progress in AI as well. The agreement should also halt efforts to develop AGI—a much easier proposition to accept when that goal is understood as creating a universal machine that replaces human labor. Neither country has strong reasons to trust the other, so the agreement must be monitored through verification techniques that do not rely on the assumption of the parties’ goodwill.
A quick and pragmatic idea would be to place auditors inside the companies developing the most advanced AI models, giving them full access to the companies’ offices, communications, and AI activities, as well as the authority to report any violations of the agreement. This sounds radical, but so did the verification efforts that helped prevent the Cold War from turning into a hot thermonuclear war.
As CIA Director John Ratcliffe said of advanced AI models this summer: “It would not be […] inappropriate to describe their capabilities as being akin to digital nuclear weapons.” This technology should be treated with that level of seriousness, not the “move fast and break things” philosophy that characterizes Silicon Valley. And right now, halting the race to build machines that replace human labor—machines that the public doesn’t want and that the industry is no longer able to control—is the only sure way to avoid disaster.
(Garrison Lovely is a freelance journalist and author of Obsolete: The AI Industry’s Trillion-Dollar Race to Replace Us – and How to Stop It)

