Benchmark: GLM 5.3 Flash nvfp4 outperforms Deepseek V4 Flash

serige · reddit · 2026-08-30

The author benchmarked Deepseek V4 Flash 0731 against GLM 5.3 Flash (nvfp4) on a 2x DGX Spark setup. Results show GLM 5.3 Flash with thinking enabled scores higher on HumanEval (97% vs 93%), and the nvfp4 quantization holds up well. The main trade-off is a 256k context limit.

Original post →

More from Infra

Infra channel →