Dev Tests NVIDIA's Deepseek V4 NVFP4: 1M Context, 96% Memory at Batch 2048
HankYeomans · x · 2026-09-05
A developer shares hours of hands-on experimentation with NVIDIA's NVFP4 build of Deepseek V4, pushing 1M-token context at 96% memory usage with a batch size of 2048. More configurations are still being validated, but the results offer a real-world reference for FP4 quantized inference at extreme context lengths.
More from Infra
- That PyTorch matmul Precision Warning Is Worth Reading After All — generativist · 2026-09-06
- Qwen3.5 9B runs fully local on a phone: reasoning, code execution, and PDF generation on-device — fuzhongkai · 2026-09-05
- What breaks when AI agents run in production? Developer explores an SRE-style control layer — Fantastic-Sleep-3352 · 2026-09-05
- Rumor: US frontier labs burn massive GPUs chasing a final 1-3% model gain — SumitGup · 2026-09-05
- Arm CEO Rene Haas: The CPU Will Never Die, Even in the Age of AI Accelerators — No Priors · 2026-09-05
- Open-Source On-Device Face Swap Hits Android: Real-Time on Hexagon NPU, 66 MB APK — Few_Caregiver8134 · 2026-09-05