Fireworks delays GLM-5.3 launch to investigate 2x推理开销
AccBalanced · x · 2026-08-30
Fireworks AI detected that open-source engines required 2x longer thinking on reasoning-heavy benchmarks (AIME & GPQA) compared to the Zai API for GLM-5.3-Flash, yielding same scores but worse token efficiency. Fearing quality issues if maxtokens were hit, they paused the launch. Investigation confirmed this discrepancy across providers. Collaborating with vllmproject and Inferact, they resolved the issue, leading to a private preview and a quick API update from Zai.
Related event: Fireworks Delays GLM-5.3-Flash Launch to Fix Overthinking(2 posts)→
More from Infra
- 40-nm Memristor Chip Turns Conductance Drift Into a Feature, Beats A100 by 50-480x — maier_ak · 2026-09-01
- 40nm Neural-Dynamics Chip Uses Conductance Drift for 2.12ms Iteration Latency — maier_ak · 2026-09-01
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01