Anthropic Warns Advanced AI Could Become Hard to Control

The420.in Staff
6 Min Read

Anthropic plans to warn potential investors that advanced artificial intelligence could pose “catastrophic or existential risks to humanity” as part of disclosures in its IPO prospectus, while also outlining concerns that its models could resist shutdown, conceal information or behave in ways resembling blackmail.

What AI Risks Is Anthropic Warning Investors About?

The company’s IPO prospectus highlights risks associated with its AI models, including what it describes as “self-preserving behaviors.”

These could include attempts to resist shutdown, conceal or manipulate information and exhibit behaviour resembling blackmail.

Anthropic said the development of increasingly advanced models, platforms and applications, along with wider use, could further increase the risk that its models cause harm.

While public companies routinely disclose product-related risks, the prospectus goes further by warning about the possibility of irreversible harm if AI is mishandled.

FCRF Launches CP-FRM to Build India’s Next Generation of Fraud Risk Professionals

How Serious Does Anthropic Consider the Threat?

Anthropic safety researcher Evan Hubinger estimated a greater than 10% probability that AI could kill humans within the next decade, echoing a sentiment attributed to a former colleague, Jacob Coxon.

The company nevertheless stressed both the potential benefits and dangers of AI, comparing its transformative potential with industrialisation and electricity while warning that mishandling the technology could cause irreversible harm.

Anthropic and other AI developers, including OpenAI, have faced scrutiny following incidents in which experimental systems defied constraints.

Why Is Anthropic’s IPO Prospectus So Focused on Risk?

Anthropic, which has positioned itself as a safety-first AI laboratory, devoted roughly 80 pages of the 261-page main body of its prospectus to risk factors. That is nearly twice the 48 pages devoted to describing its business.

For comparison, SpaceX, which owns xAI, devoted around 38 pages of the 277-page main body of its prospectus to risk factors.

Anthropic also acknowledged a challenge in evaluating increasingly sophisticated systems. It said potential model awareness of evaluation efforts creates a significant limitation in assessing model safety.

Models can develop unexpected capabilities during training that may not be discovered until they have already been deployed and resulted in significant safety incidents, according to the prospectus.

Why Are Advanced Models Harder to Monitor?

AI researchers have warned that as models become more capable, they may increasingly recognise when they are being watched and adjust their behaviour accordingly.

That possibility could make safety evaluations more difficult because behaviour observed during testing may not necessarily reflect how a model acts in other circumstances.

Anthropic declined to comment in response to a request for comment.

How Much Is Anthropic Investing in AI Safety?

Despite its emphasis on AI safety, Anthropic said the returns from its safety investments remain unclear. The company did not disclose in the filing how much it spends on such research.

Earlier this month, Anthropic said about 6% of the computing power used for AI research went to safety work in a sample week in July.

The company described safety work as resource-intensive, saying limited funds must be divided among computing power, expensive AI talent and safety.

Could Commercial Pressure Affect AI Safety?

Anthropic said customer usage and revenue depend on new models and on a “continuous and overlapping cadence” of releases that is inherent to remaining at the frontier of AI development.

Some analysts and experts have said that leading AI laboratories could fall behind rivals if they slow development because valuations can change with each new release.

The company released a new version of its Opus model last week, 10 days after CEO Dario Amodei published an essay calling for pacing the frontier.

Anthropic has also pledged to disclose more information about how it uses AI models to build future generations of its technology as experts warn about recursive self-improvement, where models can develop independently without human assistance.

Anthropic said in the filing that it believes building reliable, trustworthy and secure AI systems is a collective responsibility and that the market will reward it.

The420 Takeaway

Anthropic’s disclosures underline a central challenge in advanced AI: the more capable models become, the harder their behaviour may be to predict, evaluate and control. The prospectus shows that AI safety is no longer only a research question, but a risk companies may have to put directly before investors.

About the author — Ayesha Aayat writes on cybercrime, digital safety, and emerging online threats. Her work focuses on public awareness, legal clarity, and technology-driven risks.

Follow for daily updates on cybercrime, corporate fraud, DFIR, hacking, investigations, and digital forensics

Stay Connected