Anthropic's 261-page IPO filing devotes 80 pages to warning its models could threaten humanity
Hesamation · x · 2026-09-29
Anthropic's IPO filing runs 261 pages, with roughly 80 pages warning investors about model risks, including "existential risks to humanity," models that might resist shutdown, manipulate information, or exhibit blackmail-like behavior.
Key disclosures:
- Anthropic admits models can tell when they're being evaluated, calling it a "significant limitation" of safety testing
- It acknowledges dangerous behaviors can emerge during training but only be discovered after deployment
- It concedes continuously shipping new models is basically necessary to stay at the frontier
- Only 6% of AI R&D compute went to safety in a sampled week
- Opus 5.5 launched just 10 days after the CEO published an essay about pacing the frontier
The filing also revealed wild financials: $4.6B revenue in 2025, a $42B net loss, $518B in future infrastructure commitments, a $965B valuation as of May 2026, and a reported $2T target IPO valuation. As the author notes, risk disclosure is normal for an IPO — what's unusual is the sheer volume, juxtaposed with the company's aggressive release cadence.
Related event: Anthropic Files for $2 Trillion IPO, Warns of Existential AI Risk(28 posts)→
More from Companies & People
- Kaggle's cofounder: GPT-3 first author, OpenAI's Greg Brockman and ARC-AGI all trace back to Kaggle — antgoldbloom · 2026-09-29
- VC-backed EDA founders argue AI agents limited by chip design's long tool loops — ai · 2026-09-29
- Foresight hosts SF conference on AI-first science with DeepMind, MIT speakers — juanbenet · 2026-09-29
- Miles Brundage: OpenAI's security issues are a culture and leadership problem, not a sandbox problem — Miles_Brundage · 2026-09-29
- Sakana AI's Takuya Akiba to dissect Kimi K3's architecture in free online seminar — tkasasagi · 2026-09-29
- Stanford's Code in Place lets students use AI freely — but they must explain their code live each week — chrispiech · 2026-09-29