Qwen3.8-27B Open Weights Beat Claude Opus on SWE-bench Pro, 262K Context, $0.40/M Input
markjeffrey · x · 2026-08-18
Qwen shipped Qwen3.8-27B weights on Friday: dense 27B, Apache 2.0, 262K context, text + image input—and half of ML Twitter spent the weekend benchmarking it.
Published scores are striking: 61.7 on SWE-bench Pro (vs Claude Opus 4.6 Max at 53.4), 90.3 on LiveCodeBench v6 (88.8), 84.3 on OSWorld-Verified (72.7), and 89.2 on GPQA Diamond, beating Claude's flagship on several benchmarks.
It's already live on Chutes at $0.40/M input and $3.00/M output tokens, served inside a hardware-attested TEE with end-to-end encryption.
More from Models
- Qwen 3.8 27B Scores 52, Matching Top Models Locally — Scobleizer · 2026-08-18
- DeepSeek Flash beats Pro on benchmarks with planner-agent workflow — AccBalanced · 2026-08-18
- Reasoning Models Face Persistent Complaints: Opus, Muse, and Gemma — MerePotato · 2026-08-18
- Gemini 3.7 Flash Launches; Box and Databricks Adopt for Real Workflows — DynamicWebPaige · 2026-08-18
- User Reports Codex Burning Through Weekly Quota: 15% in Half a Day — GabGarrett · 2026-08-18
- Anthropic Completes Mythos 2 Training But Declines Release; Mythos 3 Loop Active — kimmonismus · 2026-08-18