Quantized DeepSeek V3 hits 247 tokens/s decode in just 162GB
teortaxesTex · x · 2026-08-05
A developer shared benchmark results for a quantized version of DeepSeek V3 (mistakenly referred to as V4 Flash in the original tweet), achieving a remarkable decode speed of 247 tokens/s.
Commenters noted that the model's footprint is only 162GB—physically comparable to about 3mm² of stacked 3D NAND in a high-end MicroSD, with roughly the mass of a bumblebee's brain. Despite this tiny physical footprint, it is enough to contain a high-resolution "mind" that encompasses humanity's written culture.
More from Infra
- Running MiniMax H3 Locally on RTX 4070 with SageAttention Acceleration — scooglecops · 2026-08-05
- MiniMax H3 Local Test: 5-Second Video in 50s on 96GB RTX PRO 6000 — Practical_Low29 · 2026-08-05
- Silicon Valley Hits the 'Tokenpocalypse': Microsoft Sets Budgets, Uber Burns Annual Quota — 量子位 · 2026-08-05
- Opus Crammed into a Single Device in 7 Months: Edge AI Explosion Accelerates — teortaxesTex · 2026-08-05
- SpaceX Revenue Nearly Doubles to $7.8B, AI Compute Contracts Surge 250% — rohanpaul_ai · 2026-08-05
- MiniMax H3 Local Test: RTX PRO 6000 Outperforms B300 by Almost Half in Inference Time — Resident_Sympathy_60 · 2026-08-05