Swift 1.5 matches Qwen3.8 27B quality on M5 Max while writing 34% fewer tokens
DerTomsn · reddit · 2026-09-30
A Redditor benchmarked Swift-1.5-Qwen3.8-27b (oQ8e, MLX) on an Apple M5 Max 64GB via llm-bench.io with 262k context and thinking at xhigh, against base Qwen3.8 27B.
- Token output: Swift averages 51k generated tokens per full run vs 77k for Qwen3.8 — biggest savings in code generation (28.6k vs 40.1k) and research (11.1k vs 21.2k); about 3/4 of Swift's output is still reasoning, just less of it
- Speed & quality: 34.6 vs 32.8 tok/s generation, 400 tok/s prompt processing for both, and LLM-judge scores of 85.8 vs 85.0 — a tie within run-to-run variance (84.3–87.2)
- Wall time: full benchmark averaged 24 min for Swift vs 38 min for Qwen3.8
Bottom line: equal quality, roughly a third faster thanks to leaner reasoning. The author plans to use it as a daily driver. All eight raw benchmark runs are linked.
More from Infra
- AWS Bedrock rolls out GPT-6 Sol, GPT-6 Luna and Claude Opus 5.5 — emmanuelvivier · 2026-09-30
- NVIDIA releases Nemotron Labs for running Nemotron models locally on DGX Spark & Station — NVIDIAAI · 2026-09-30
- AI infra will look like a distributed energy grid, not a few giant plants — sarahookr · 2026-09-30
- Efficient Computer Raises $97M to Scale Chips Claiming 10x Energy Efficiency for AI — rebeccakaden · 2026-09-30
- UK town fights 850-acre AI datacentre plan, rebuts Osborne's nimby charge — nordicinst · 2026-09-30
- CSET Report: AI Chip Smuggling Undermines Export Controls, Can Location Verification Help? — chrisrohlf · 2026-09-30