Positron Valued at $5B in 7 Months, Betting AI Inference Chips Don't Need HBM

快鲤鱼 · wechat · 2026-09-14

AI inference chip startup Positron raised $875M at a $5B post-money valuation — 5x its prior round just 7 months earlier — led by NEA and Jim Clark, with Qatar Investment Authority, Cisco Investments and SemiAnalysis founder Dylan Patel participating. Its core thesis: the real bottleneck in LLM inference is memory bandwidth, not compute. Its next-gen Asimov chip ditches expensive HBM for commodity LPDDR5X, using a systolic-array architecture with local memory per compute block that it claims achieves >90% memory bandwidth utilization (vs 10–30% for typical GPUs), up to 2.3TB memory per chip for million-token contexts. Simulations show up to 26x GB300 NVL72 in tokens per dollar — but Asimov hasn't taped out yet (planned late 2026, production H2 2027).

The team is operator-heavy: CTO Thomas Sohmers (Thiel Fellow, ex-Lambda/Groq) and a CEO who scaled Lambda's GPU cloud past $500M ARR. The current Atlas product (still HBM-based) runs on 50+ racks at Oracle Cloud. The piece also compares Groq, Etched, SambaNova and Cerebras, arguing they all bet that inference costs will eventually exceed training costs, and that tokens-per-dollar will replace peak FLOPS as the metric that matters.

Original post →

More from Venture

Venture channel →