DeepSeek V4 Flash Hits 32.7 tok/s on AMD Strix Halo via 98GB GGUF
tensorqt · x · 2026-08-12
Lucebox collaborated with Geometric to deploy DeepSeek-V4-Flash-0731 on a 128 GB AMD Strix Halo system.
- Local Deployment: The model is quantized into a single 98.29 GB GGUF file, loading unsplit on the Radeon 8060S iGPU without sidecar files.
- Quality: It scores 82/92 on the full evaluation and a perfect 17/17 on COMPSEC-17, maintaining near-lossless quality.
- Speed: Decode speed reaches 18.1 tok/s in quality mode. Using the faster DSpark helper mode pushes speeds up to 22.3–32.7 tok/s at a slight cost to accuracy.
Related event: AMD Strix Halo Runs DeepSeek V4 Flash at 32 tok/s(2 posts)→
More from Infra
- Running MiniMax H3 on Low VRAM Fried My GPU, Beware — ROBOTTTTT13 · 2026-08-12
- StarCloud Explores Space Data Centers: Launching AI Hardware into Orbit — DavidLinthicum · 2026-08-12
- gakonst adds confidential compute support to nanocodex for verifiable confidential AI — AccBalanced · 2026-08-12
- Alibaba Cloud's CUBE 5.0 Modular Design Builds AI Data Centers in 100 Days at 10% Lower Cost — rohanpaul_ai · 2026-08-12
- CoreWeave Secures $2.6B Credit Facility Backed by Long-Term NVIDIA GPU Value — OnlineInference · 2026-08-12
- Poolside's Laguna Model Acceleration Challenge Yields 2.6x Speed Boost by Community — gajesh · 2026-08-12