Benchmark: GLM 5.3 Flash nvfp4 outperforms Deepseek V4 Flash
serige · reddit · 2026-08-30
The author benchmarked Deepseek V4 Flash 0731 against GLM 5.3 Flash (nvfp4) on a 2x DGX Spark setup. Results show GLM 5.3 Flash with thinking enabled scores higher on HumanEval (97% vs 93%), and the nvfp4 quantization holds up well. The main trade-off is a 256k context limit.
More from Infra
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01
- Samsung shifts to 8-layer HBM4E for Nvidia with ~20% higher speed spec — 创业邦 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01