Dev Tests NVIDIA's Deepseek V4 NVFP4: 1M Context, 96% Memory at Batch 2048

HankYeomans · x · 2026-09-05

A developer shares hours of hands-on experimentation with NVIDIA's NVFP4 build of Deepseek V4, pushing 1M-token context at 96% memory usage with a batch size of 2048. More configurations are still being validated, but the results offer a real-world reference for FP4 quantized inference at extreme context lengths.

Original post →

More from Infra

Infra channel →