Qwen3.8-Flash-Next on a 5090 beats Claude Code on airbench: 11 min vs 14 min
dh7net · reddit · 2026-10-11
A Reddit user ran the IQ3S quant of Qwen3.8-Flash-Next via Strata (local inference server) with the pi coding agent on a single RTX 5090 + 64GB RAM, scoring 100% on the airbench everyday-task suite in 11 minutes; Claude Code got the same score in 14 minutes.
Reproducible setup details:
- Hardware: GPU with ≥24GB VRAM, 64GB RAM, 90GB disk. RAM is the real limit—the IQ3S quant keeps 50GB of experts in RAM; Strata estimates 78GB for 256k context.
- Model server: Strata downloads 84GB on first start; Docker build with CUDAARCHITECTURES=120 (5090) or 86 (3090); run flags include --expert-cache auto --prefill auto --spec 4 --kv int8 --vision; serves an OpenAI-compatible API on port 9000.
- pi agent: models. sets up the local provider; settings. raises output cap to 32k and compaction reserve to 64k (default 16k lets one image result overflow the window).
- The author wrote a small proxy that shrinks oversized output requests (when ≥8k tokens remain) and rewords overflow errors so pi compacts automatically.
- Same setup comparison: omp 18.4.2 scored 48/49 in 16 minutes.
Full commands and configs included—copy-paste ready for local open-model coding agents.
More from coding & agent
- Turingo detects AI writing by replaying document revision history, not text predictions — sethlazar · 2026-10-11
- Claude is an underrated used-car hunting tool: market modeling, scraping, daily alerts — eherrerosj · 2026-10-11
- AI agent tunes Triton kernels on AMD MI210, flipping grid order yields 1.29x speedup — zmkzmkz · 2026-10-11
- OpenAI's Decisions API enters public beta, now callable via Apple's Foundation Model framework — rxwei · 2026-10-11
- Codex vs Claude Code: the quota reset economics and the optimal time to press — juntao · 2026-10-11
- Astra in Codex browses for connector specs, models parts and drops them into assemblies — burhop · 2026-10-11