Zach Mueller to AIPerf-benchmark GLM 5.3 Flash, DeepSeek v4 Flash, Qwen Flash Next on x8 PCIe5 GPUs
TheZachMueller · x · 2026-10-09
Zach Mueller is collecting requests for an AIPerf comparison this weekend, running models on x8 PCIe5 across 4 Pro vs 4 Max-Q hardware, preferring stable NVFP4 quants as the low end. Current candidates: GLM 5.3 Flash, DeepSeek v4 Flash, and Qwen Flash Next. He says the goal is to publish useful baselines ahead of the backplane release.
Related event: Hugging Face Engineer to Benchmark Flash Models on NVFP4 with AIPerf(2 posts)→
More from Infra
- Synopsys eyes Chinese AI labs for chip design, forecasts $11.15bn FY27 revenue — pstAsiatech · 2026-10-09
- TRL v1.15 defaults to fused LM head, extending training sequences up to 6.9x — LysandreJik · 2026-10-09
- NVIDIA NeMo-DCR cuts 1T-model RL weight sync from 87.5 min to 150 sec — IanAndrewsDC · 2026-10-09
- Deno team joins Cloudflare to merge with Workers and build the default server programming model — irvinebroque · 2026-10-09
- China built solar and wind at colossal scale; next step is powering the AI race — pstAsiatech · 2026-10-09
- Hugging Face ships inference conceptual guide on prefill, decode, and KV cache — mervenoyann · 2026-10-09