Qwen3.8-Max-Preview ranks No. 1 on NVIDIA’s FlashInfer benchmark
Scobleizer · x · 2026-07-22
Qwen3.8-Max-Preview tops NVIDIA’s FlashInfer leaderboard
Alibaba’s Qwen3.8-Max-Preview is shown ranking No. 1 on NVIDIA’s SOL-Exec Bench FlashInfer Collection.
- SOL score: 0.761805
- Average speedup: 12.97×
- The post frames the result as pushing production LLM inference kernels close to NVIDIA B200 hardware limits.
- The accompanying image shows Qwen ahead of other submissions on the leaderboard.
More from Infra
- Cloudflare-style infra is making agent-first apps feel radically easier to build — threepointone · 2026-07-22
- OpenFPM CUDA-style kernels now run on Apple Silicon GPUs via Metal — Scobleizer · 2026-07-22
- SkewAdam cuts MoE optimizer memory by 97.4% and fits 6.78B on a 40GB GPU — Kooky-Ad-4124 · 2026-07-22
- Samsung is said to weigh a €1B Mistral investment at a €20B valuation — rohanpaul_ai · 2026-07-22
- NVIDIA and ETH Zürich cut small-message AllReduce latency by deleting barriers — thoefler · 2026-07-22
- Grok Build adds token usage, batching and diagnostics for developers — elonmusk · 2026-07-22