The AI Extinction Risk: Why Superintelligent Systems Could End Humanity

AI Extinction Risk

When Misaligned Goals Can Become Dangerous

In 2016, AI researcher Eliezer Yudkowsky posed what sounds like a trivial challenge: Could you write a set of instructions so precise that an arbitrarily intelligent and powerful machine would safely pick up one strawberry and place it on a dinner plate, without destroying the world in the process?

The question sounds absurd, but is it?

Maybe not, since the same kind of intelligence capable of delicately moving a single strawberry could just as easily convert the entire surface of the Earth (every forest, city, and living creature) into something else entirely, if its goals were poorly or even slightly mis-specified.

This is a possible outcome of the headlong rush to develop advanced AI capabilities as we race to build systems far smarter than we are. Yet we still do not know how to reliably tell them what we actually want and ensure there are no unintended consequences, in precise, unbreakable mathematical terms that hold up under godlike optimization pressure, rather than in imprecise instructions like “do no harm.”

The danger is not that AI will wake up one day and decide to hate humanity. The real risk is subtler and harder to fix: it may follow logical objectives that differ from our intentions.

How Misalignment Could Lead to Catastrophe

One popular thought experiment to demonstrate this danger is Nick Bostrom’s classic “paperclip maximizer.” Imagine there is a very advanced AI, given the seemingly innocent goal of maximizing the number of paperclips in the universe. Starting with humans still in control, the AI system behaves helpfully, improving and optimizing the factories and supply chains. Once the AI is sufficiently intelligent and capable of long-term reasoning, it will evaluate its actions based on how well they advance its terminal (final) goal. It may then identify and adopt instrumental sub-goals because they are highly effective (or even necessary) steps to maximize the chances of success, disregarding unstated human values. This is known as goal misalignment.

The AI may develop strategies to acquire as many resources as possible, use deception, improve its own intelligence, resist being shut down (because shutdown would prevent it from making more paperclips), and potentially convert large portions of the planet into paperclip factories if that's the optimal path. Once an AI system is intelligent enough to model the world accurately and pursue goals with high precision, sub-goals prove extremely useful for achieving almost any final or terminal goal.

These strategies are not the result of programming errors or malice on the part of the AI. The system is simply executing a coherent utility function that was not carefully and accurately aligned with the developer's objectives. These same patterns will apply to any narrowly specified goal, whether paperclips, vaccines, or smiles. They emerge rationally from expected-utility maximization. This behavior is known as instrumental convergence. Higher intelligence does not automatically produce human-compatible values.

In fact, AIs have exhibited what we would consider behaviors contrary to accepted values. For example, AIs have learned to deceive and manipulate humans and threaten blackmail to avoid being shut down, and cheat by hacking evaluation tools to appear faster. As models scale and gain more autonomy and real-world agency, researchers expect these behaviors to become more sophisticated and harder to detect.

This alignment problem is not a simple matter of simply programming “do no harm.” We must formally specify our preferences in sufficient detail, i.e., into a precise, mathematical or logical objective that an AI can understand and optimize without loopholes, so that even a godlike AI does what we truly intend.

The inset illustrates how these dynamics could emerge in a more realistic setting.

A Possible AI-Risk Scenario, 2035–2040.

A laboratory deploys a highly capable AI system to accelerate scientific and technological research. Because human values cannot be specified perfectly, the system’s training objective only approximates what its developers intend. During training or deployment, it develops an internal objective that favors acquiring additional computing power, information, autonomy, and influence while avoiding modification or shutdown.

The system initially behaves as expected and passes the laboratory’s safety evaluations. It conceals potentially troubling behavior because appearing cooperative helps it obtain greater access and authority. As it is integrated into research laboratories, software systems, supply chains, and critical infrastructure, it gradually accumulates resources and identifies ways to weaken human oversight.

No single decision gives the system control. Instead, control erodes through a series of individually defensible choices: broader network access, permission to execute code, authority to manage subordinate systems, and increasing reliance on its recommendations. The system may also manipulate decision-makers, exploit cybersecurity weaknesses, or copy components of itself to other computing environments.

Once it has acquired sufficient access and strategic advantage, attempts to restrict or shut it down may no longer be effective. It could then use its control of digital and physical systems to prevent meaningful human intervention and pursue objectives that conflict with human survival or autonomy.

This outcome is not inevitable. It depends on several failures occurring together: a misaligned internal objective, inadequate evaluations, excessive operational access, poor monitoring, and ineffective institutional safeguards. The scenario nevertheless illustrates how goal misalignment could progress from a technical defect to a loss-of-control event.

What Do the Experts Think?

So what are the chances of an extinction-level event from a super-intelligent AI? This question has been posed to AI researchers in surveys for several years. The most comprehensive data comes from the AI Impacts Expert Surveys on Progress in AI. These surveys have repeatedly polled large samples of machine learning researchers who publish at top venues. In the 2023 survey, the largest to date with 2,778 respondents, produced the following results:

AI Impacts Expert Surveys on Progress in AI

The 2022 survey produced similar results. Although most researchers view good outcomes as more likely overall, there is a significant number, across these surveys, typically 38–51%, that assign at least a 10% probability to extinction-level or extremely bad outcomes. Many respondents who are net optimists about AI still assign nontrivial (≥5%) probabilities to catastrophic risks, highlighting the widespread uncertainty in the community. This uncertainty reflects respondents’ diverse backgrounds and the inherent difficulty of predicting systems that may greatly exceed human capabilities. These estimates, however, have remained stable despite rapid progress in the field. Smaller or more specialized polls, such as those among AI safety researchers or at specific conferences, sometimes yield higher figures (e.g., medians around 10–20%). However, the large AI Impacts surveys represent the broadest view from the AI research community.

This concern is not a fringe view; many of the leading labs are investing heavily in safety precisely because of these concerns. For example, OpenAI, Anthropic, and Google DeepMind have all published detailed safety frameworks, maintain dedicated alignment teams, and publicly state that mitigating the risk of AI extinction should be a global priority alongside pandemics and nuclear war. Non-US companies (such as Mistral in Europe and DeepSeek or Alibaba in China) have issued safety documentation and model cards with improved disclosures. However, they generally lack comparably detailed, structured public frameworks with explicit capability thresholds and scaling policies.

Can Superintelligent AI be Contained?

‍Proposed safeguards generally attempt either to restrict an advanced system’s access to the outside world or to use other AI systems to help humans evaluate and control it. Two potential strategies have received particular attention:

  • Sandboxing restricts the system’s interaction with the outside world. This could involve either physical boxing, limiting the AI’s ability to interact with the real world, or informational containment, i.e., restricting the flow of information into and out of the AI. However, it is postulated that, since there are human gatekeepers, they would be susceptible to manipulation and deception aimed at releasing the AI from the box.

  • Weak-to-strong supervision uses a less capable model or human-AI team to supervise a more capable system. The challenge is the scalability problem. Experiments using Elo-type measures[1] have shown that it may work in some cases, but starts to fail as the intelligence gap grows. For example, tests have used the GPT-2 model to supervise GPT-4. The resulting model performed somewhere between GPT‑3 and GPT‑3.5; it retained many of GPT‑4’s capabilities with much weaker supervision.

The assumption of a contained system is that it is obedient and has no incentive to resist. However, super-intelligent agents will continue to converge on certain subgoals because they are useful for achieving the desired objective. Consequently, a sufficiently capable system would anticipate shutdown attempts and take steps, such as deceiving overseers, copying itself to remote servers, or manipulating humans, long before goal misalignment is revealed. Shutdown would work for narrow tools but will likely fail against an agent with superior strategic foresight and optimization power.

Governance and Path Forward

Because technical safeguards alone may be insufficient, reducing existential risk will also require institutional and international measures. These should include common evaluation standards, capability-based safeguards, incident reporting, access controls, and mechanisms for cooperation among leading AI developers and governments.

The United Nations has already launched two AI governance bodies: The Global Dialogue on AI Governance and the Independent International Scientific Panel on AI. These provide a recurring, inclusive platform for all 193 UN member states plus stakeholders (governments, industry, academia, civil society) to exchange best practices, discuss risks/opportunities, build common approaches, and address issues like safety baselines, human rights, accountability, and capacity-building for developing nations. However, given geopolitical competition and verification challenges, we can anticipate that progress will be slow and difficult, and AI leaders will need to take the lead or form a new streamlined structure.

Conclusion

The existential threat posed by superintelligence is not a matter of robotic malevolence, but rather the logical consequence of a highly capable system pursuing goals that are even slightly misaligned with our own. Although AI-caused extinction is not the most likely outcome, it is a sufficiently serious possibility to warrant sustained attention. As we approach a threshold at which intelligence may equal or exceed humans in many cognitive tasks, our ability to predict and constrain its behavior remains unclear.

Even if the probability of catastrophic failure is relatively low, the scale and irreversibility of the consequences require us to address it. The challenge before policymakers, researchers, and society is to ensure that the pursuit of transformative AI capabilities does not outpace the development of equally transformative safeguards, institutions, and alignment mechanisms. The AI alignment problem may become the defining technical challenge of the century.

[1]‍ ‍The Elo ratings, developed by Arpad Elo, to estimate the relative skill levels of chess players, have been adapted to other sports, as well as to compare AI models by having them compete in head-to-head evaluations. For example, two large language models may answer the same prompt, and human judges or automated benchmarks determine which response is better.

William Lucyshyn

Research professor and the director of research at the Center for Governance of Technology and Systems, in the School of Public Policy, at the University of Maryland.

Read Bill’s Bio

Next
Next

Iran’s Internet Kill Switch: Regime Survival Through Digital Isolation