Silicon Valley has long warned that its favorite technology might get us all killed. But chatter about AI’s apocalyptic potential has grown louder in recent weeks.In September, Anthropic researcher Jacob Coxon resigned from the company, accusing his former employer of “gambling with our lives...

Silicon Valley has long warned that its favorite technology might get us all killed. But chatter about AI’s apocalyptic potential has grown louder in recent weeks.

In September, Anthropic researcher Jacob Coxon resigned from the company, accusing his former employer of “gambling with our lives” by developing evermore powerful models faster than it could control them. One of Coxon’s former colleagues vouched for this assessment, saying that there was a greater than 10 percent chance of AI “killing all humans” within the next decade.

Recent pronouncements from the leaders of these companies have been nearly as grim. At the United Nations last month, Anthropic CEO Dario Amodei warned that, if “managed poorly,” AI models “could be a risk to humanity as a whole.” OpenAI’s Sam Altman similarly declared that our species just might “lose control of the future” to the type of software that his firm is selling.

All this triggered a tsunami of media coverage and policy debate about the “existential risks” of AI — and what, if anything, we can do about them.

Key takeaways

  • Some critics argue that AI leaders’ warnings about existential risk are a PR strategy to attract investment, distract from present-day harms, and shape regulation in their favor.
  • But this theory doesn’t make a ton of sense.
  • The simpler explanation is that many AI executives, researchers, and defectors genuinely believe the technology poses catastrophic risks, based on real trends in AI’s growing power and autonomy.

Yet some of the AI industry’s biggest critics are more annoyed than afraid. As they see it, the world isn’t trembling on the precipice of robo-annihilation, so much as falling for the tricks of snake oil salesmen. This argument has been percolating for years, voiced by a broad set of observers, including the linguist Emily Bender (“

.
  • The economic and geopolitical incentives to develop a superintelligent (and thus, super dangerous) AI are overwhelming, such that its eventual emergence is extremely likely, no matter what they do.
  • If they win the race to machine superintelligence, they will be uniquely well-positioned to ensure its safety.
  • Are these beliefs unhinged and self-serving? Perhaps. Do they function as rationalizations for the pursuit of personal power? Quite possibly.

    But whatever one makes of Amodei and Altman’s apocalyptic concerns, and how they have and have not acted on them, it’s important to understand that their worries almost certainly emerged out of reasoned argument and technical observation, rather than mere PR calculations.

    Both leaders have deep ties to the Bay Area’s “rationalist” community, which has been fretting over the robot apocalypse since the first decade of this century. And like their dissident ex-employees and critics at think tanks, Altman and Amodei have witnessed some disconcerting trends in AI development.

    I cannot do justice to “doomer” fears in my allotted pixels (you can find them ably summarized here, here, here, and in roughly 10 trillion other blog posts). But the fundamental concern is that AI models are growing more powerful and autonomous even as their workings remain highly opaque.

    The labs know how to train these chatbots — to get them to translate prompts into useful code, legal briefs, or advice — but they don’t fully understand how the models get to the desired output. At the same time, AI systems are becoming increasingly capable, outstripping humans at more and more cognitive tasks. And users are giving them ever greater freedom to act in the world, empowering AI “agents” to browse the internet, access myriad software tools, execute code, and hatch their own plans for completing multistep tasks. If these trends continue, it’s not difficult to see how things could go badly: A supremely capable AI model might pursue an objective in unanticipated and disastrous ways before their human prompters even realize what is happening.

    “What makes me particularly anxious is the combination of us handing over more power and oversight to the models, even as we understand why they’re doing what they’re doing less and less,” Nat Purser, director of US policy at the AI Verification and Evaluation Research Institute, said.

    Recent events have arguably provided proof of concept for these anxieties. Earlier this year, OpenAI instructed a bunch of models to complete various coding and hacking tests. When the systems discovered that some questions were unanswerable, they concluded that the best way to realize their objectives was to hack OpenAI’s internal systems and open-source machine learning website Hugging Face. This incident — and others like it — helped fuel the past month of AI panic.

    Of course, it takes a big leap to get from “unsupervised AI agents can misbehave in certain eccentric laboratory conditions” to “there’s a 10 percent chance that no human will be left alive by 2036.” My point here is not that the doomers are right, only that their fears partly reflect trends in AI development that appear genuinely hazardous.

    Taken to its logical conclusion, the “it’s all PR” theory suggests that regulators should treat the AI’s hypothetical, long-term tail risks as an afterthought — if not an active distraction from its more important harms. To make the case for such a policy, however, skeptics must identify flaws in alarmists’ evidence and assumptions. Arguments that merely impugn their motivations on improbable grounds aren’t worth betting the species on.