OpenAI's own inference chip, Nvidia's ~$12.9B Hugging Face bid, and cheap open Qwen3.8-Flash
ksraj1001 · reddit · 2026-08-29
The author ties three stories into one theme — inference economics is now the main event:
- OpenAI's first custom inference chip 'Jalapeño' (with Broadcom, Celestica; Samsung on HBM4): claims 1.5–1.9× throughput per kW and 1.7–3.6× lower end-to-end latency vs Nvidia GB200/GB300 racks, deploying end of 2026 — vendor-reported benchmarks, so grain of salt.
- Nvidia closing in on acquiring Hugging Face for $12.9B: HF is the de facto neutral hub for open models; Nvidia owning it raises neutrality and hardware-default questions.
- Alibaba's Qwen3.8-Flash: 125B open weights, reportedly competitive with Opus 4.6 and DeepSeek V4-Flash at aggressive pricing; Qwen passed 3B downloads, ahead of Meta and Google, with revenue-sharing tests for large commercial users.
Context: this month's biggest rounds were inference infra (Fireworks AI $1.5B, Together AI $800M), while Claude had a rough uptime month. Practical advice: stop treating your model provider as fixed — benchmark cheap open models on your real workload and build fallbacks — while watching the concentration risk as chips, the open-source hub and frontier models consolidate.
More from Venture
- VC on AI-era growth: 10-15% monthly compounding for 4-6 years beats 50x-or-nothing — JosephJacks_ · 2026-08-30
- Broken model wrappers sell Seedance 2.5 at $0.02/sec, 10x below standard pricing — oyacaro · 2026-08-30
- a16z's Vijay Pande: Biology shifting from discovery to engineering — TechCrunch AI · 2026-08-30
- Early-Stage Venture Alpha Lies in Independent Action, Not Over-Intellectualization — arian_ghashghai · 2026-08-30
- Publishing to the Claude Connector Directory: 60 signups in 2 days, 3 sales, then the discoverability cliff — kkomelin · 2026-08-30
- Goldman's AI-free S&P 500 index now beats the real one; AI names were 45% of weight — rohanpaul_ai · 2026-08-30