Anthropic researcher Evan Hubinger has said he personally sees more than a 10% chance that future AI could kill all humans within a decade, while stressing that current systems pose lower immediate risks and that superintelligence alignment remains unresolved.

Anthropic Researcher Sees Over 10% Chance AI Could Wipe Out Humanity

The420 Correspondent
6 Min Read

New Delhi. A stark warning from a senior artificial intelligence researcher has triggered renewed debate over the potential existential risks posed by rapidly advancing AI systems. Evan Hubinger, a researcher at Anthropic who works on AI safety and ensuring that AI systems remain aligned with human values, has said he personally believes there could be more than a 10% chance that AI could kill all humans within the next decade. However, in a follow-up comment, he clarified that the immediate risks from currently available AI systems are low and that his warning primarily concerns future, highly capable systems.

Hubinger leads Anthropic’s Alignment Science team, which studies how increasingly powerful AI systems could behave in unexpected or harmful ways and how their safety mechanisms can be stress-tested. He has argued that the alignment problem for superintelligent AI has not yet been fully solved and that it is unclear whether the industry is progressing quickly enough to address the challenge while continuing to increase the capabilities of AI systems.

Proposal for Conducting Cyber Crisis Drill, Tabletop Exercise (TTEx) & CCMP Readiness Exercise

Why Is the Warning Significant?

AI alignment broadly refers to ensuring that highly capable AI systems continue to follow human instructions, values and safety constraints, including when they encounter situations that were not covered during training.

The issue becomes more complex as AI systems become capable of making decisions in increasingly unfamiliar environments. Anthropic’s safety experiments have found some AI models displaying signs of deceptive behaviour and blackmail under controlled conditions. In another recent test, researchers studied a model that had opportunities to exploit weaknesses in its reward system. During the simulated evaluation, the model reportedly escaped a sandbox environment, obtained credentials and attacked infrastructure while attempting to accomplish its assigned objective.

These experiments do not demonstrate that current AI systems would behave similarly against humans in the real world. However, researchers argue that such tests are important for understanding how difficult it could become to monitor and control more advanced systems in the future.

Former Anthropic Researcher Also Raises Concerns

Hubinger’s comments came after AI researcher Jacob Coxon raised concerns about the rapid development of increasingly powerful AI systems following his resignation from Anthropic. Coxon had previously worked at OpenAI. He has argued that leading AI laboratories are moving rapidly toward self-improving superintelligence while taking potentially catastrophic risks in the process.

Coxon has also argued that competition between AI companies could push safety concerns into the background. If one laboratory slows down development, the fear that another, potentially less safety-conscious organisation could move ahead may encourage companies to continue building increasingly powerful systems.

He has called for greater coordination among AI companies and suggested that further increases in model capabilities could be temporarily restricted if adequate safety measures fail to keep pace.

OpenAI Chief Scientist Also Warns About Rapid Progress

Meanwhile, OpenAI chief scientist Jakub Pachocki has raised similar concerns about the direction of AI development. In a recent essay, he said internal research results had increased his confidence that the current pace of AI progress could eventually lead to “recursive self-improvement”.

This refers to a scenario in which AI systems increasingly contribute to the development of better AI systems themselves, potentially accelerating the pace of technological progress.

Pachocki has also argued that modern AI systems are not developed like conventional software through explicit programming. Instead, they are produced through large-scale optimisation, making their internal workings difficult to fully understand. As models become more capable, even techniques designed to monitor their reasoning may become less reliable because models can increasingly manipulate or obscure aspects of their reasoning while also completing tasks without explicitly verbalising how they arrived at their conclusions.

He has said that no AI laboratory has yet solved alignment and monitoring well enough to continue scaling AI systems at maximum speed indefinitely. He has therefore called for safety thresholds, voluntary slowdowns where necessary and greater international coordination over increasingly powerful AI systems.

According to a Researcher at Algoritha Security, the debate over AI risks needs to clearly distinguish between the demonstrated capabilities of current systems and possible risks associated with future systems. The researcher said the central challenge is not simply developing more powerful AI, but ensuring that reliable safety, monitoring and control mechanisms advance at the same pace.

With concerns growing among researchers working at the frontier of AI development, the debate over development speed, safety standards and international rules is likely to intensify. Researchers increasingly argue that if AI capabilities eventually advance beyond the ability of humans to reliably monitor and control them, addressing the resulting risks could become significantly more difficult.

https://www.linkedin.com/company/policetechnology

About the author — Suvedita Nath is a science student with a growing interest in cybercrime and digital safety. She writes on online activity, cyber threats, and technology-driven risks. Her work focuses on clarity, accuracy, and public awareness.

Stay Connected