antirez Optimizes DwarfStar: 170 t/s Generation and 22k tokens/s Prefill on Station
antirez · x · 2026-08-17
antirez was asked about DwarfStar's performance on Station with MXFP4 and the Flash model instead of PRO. After about 48 hours of optimization, it achieves 170 t/s generation without DFlash and 22k tokens/s prefill, with potential for further improvement.
Related event: antirez Tunes DwarfStar to 170 t/s Generation on Station(2 posts)→
More from Infra
- Bittensor Subnet 118 Adds Ultra-Cheap Inference, Joining Major AI Providers — markjeffrey · 2026-08-17
- Meta to rely on Nvidia Blackwell, AMD Helios in 2026, accelerate custom MTIA in 2027 — Beth_Kindig · 2026-08-17
- Stripe to Acquire OpenRouter for Over $7B, 5.4x May Valuation — rohanpaul_ai · 2026-08-17
- Wici One claims to solve local VRAM limits via NVMe offloading — Torodaddy · 2026-08-17
- Qwen3.8-27B hits 206 tok/s on single RTX 5090 via SGLang — StefanoGogioso · 2026-08-17
- CoreWeave Revenue Outpaces Big Cloud Early Stages Amid AI Pivot — a16z · 2026-08-17