Qwen3.8 preview tops NVIDIA’s CUDA kernel optimization benchmark
智东西 · wechat · 2026-07-22
Qwen3.8 preview tops NVIDIA’s kernel-optimization benchmark
Alibaba’s Qwen3.8 preview reportedly ranked first on NVIDIA’s SOL-ExecBench FlashInfer-Bench subset, beating Tencent Hunyuan, human developers, and Cursor. The benchmark is designed to measure how well models and agents optimize CUDA kernels on Blackwell B200 hardware, with scores tied to proximity to the hardware’s theoretical speed-of-light (SOL) limit.
- Bench scope: 235 CUDA kernel optimization tasks, extracted from 124 models and spanning LLMs, diffusion, vision, audio, video, and hybrid workloads.
- What the score means: higher SOL scores indicate kernel implementations that get closer to peak compute-and-memory efficiency, not just faster than a software baseline.
- Why it matters: NVIDIA says kernel optimization is a rare real-world task with clear feedback, making it useful for training models that can iteratively improve from hardware execution signals.
- Qwen3.8 details: the preview reportedly won the FlashInfer-Bench subset, which focuses on inference kernels such as fused attention, FP8 MoE, and RMSNorm, and the team says the model ran through tens of thousands of tool-call iterations during evaluation.
The article argues that strong performance here suggests broader potential for closed-loop self-improvement in systems work, compiler-like reasoning, and hardware-aware optimization.
Related event: Alibaba's Qwen3.8-Max-Preview Tops NVIDIA FlashInfer Benchmark(3 posts)→
More from Infra
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11