Mistral Large 4 trained on just 4,000 Grace Blackwell GPUs vs 100,000 for Astra
steipete · x · 2026-10-07
NVIDIA revealed that Mistral Large 4, now in public preview, was trained on nearly 4,000 Grace Blackwell Superchips billed as "frontier scale, built in Europe." A repost highlights the stark disparity: the competing Astra model was trained on 100,000 GPUs — 25x more. Whether Mistral's compute-efficient approach can keep pace is the open question.
Related event: Mistral Large 4 Trained on ~4,000 Grace Blackwell Chips in Europe(4 posts)→
More from Infra
- 21M model + 6.4B SSD-resident lookup table matches a 114M dense model — fechyyy · 2026-10-07
- Ollama 0.35 adds local decision models from Cloudflare, Together and Bespoke — Technovangelist · 2026-10-07
- CostGraph becomes a drop-in InfraCost replacement for comparing GPU prices from L40 to A100 — saheedniyi_02 · 2026-10-07
- Oki Home launches a $1,799 Memory Computer running a 27B model locally with up to 16TB Memchip storage — Scobleizer · 2026-10-07
- AveniatsHub runs Qwen3 and Gemma3 fully offline on Android — MiraMooreIloveLLM · 2026-10-07
- Phonon-2 hits 606x real-time on a MacBook Air: one hour of speech in 6 seconds — julianweisser · 2026-10-07