Google search engine


AI company flags potential risks including models resisting shutdown, concealing information and exhibiting behaviour resembling blackmail

Anthropic has warned potential investors that increasingly advanced artificial intelligence could pose “catastrophic or existential risks to humanity”, highlighting the potential for AI models to develop unexpected and harmful behaviours, Reuters reported.

The warning was contained in Anthropic’s IPO prospectus, which outlines risks associated with the technology as the company prepares to go public.

According to Reuters, Anthropic said its AI models could exhibit “self-preserving behaviours”, including attempts to resist shutdown, conceal or manipulate information and behaviour resembling blackmail.

“Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm,” the company said in the filing.

Anthropic, the creator of Claude AI models, has positioned itself as a safety-focused AI company while competing with major developers, including OpenAI.

Risk factors dominate IPO filing

Anthropic devoted around 80 pages of its 261-page main prospectus to risk factors, nearly twice the 48 pages dedicated to describing its business, Reuters reported.

The company warned that AI models can develop unexpected capabilities during training that may remain undetected until after deployment, potentially resulting in significant safety incidents.

“Potential model awareness of our evaluation efforts creates a significant limitation on our ability to assess model safety,” Anthropic said.

businessMore from Business

The company also warned that increasingly capable models may recognise when they are being monitored and adjust their behaviour, making safety assessments more difficult.

Safety investment faces trade-offs

Anthropic said AI safety research is resource-intensive and competes with other major expenses, including computing power and AI talent.

The company has not disclosed its total spending on safety research. However, it said earlier this month that about 6 per cent of the computing power used for AI research in a sample week in July was devoted to safety work, Reuters reported.

At the same time, Anthropic said revenue depends on the development and adoption of new models, making frequent releases important to maintaining its position at the frontier of AI.

The company recently released a new version of its Opus model, shortly after CEO Dario Amodei called for greater caution in the development of frontier AI.

Anthropic said it believes building reliable, trustworthy and secure AI systems is a collective responsibility and that the market will reward such systems, according to Reuters.

Google search engine