Qwen3.8-27B runs at 63 tok/s on Mac Studio
remilouf · x · 2026-08-17
A user shared a benchmark of running the Qwen3.8-27B model locally on a Mac Studio, achieving an inference speed of 63 tokens per second, noting that it is very usable.
More from Infra
- Managing the software stack around local LLMs — IllegalStateExcept · 2026-08-18
- Frontier Intelligence Gets as Cheap as Cloud Storage, Shifting Enterprise AI Race to Routing and Orchestration — krishnan · 2026-08-18
- Investing in the Machine Economy: Scarce Assets Beyond AI Generation — 0xSammy · 2026-08-17
- Investigation: Rare Books Tracked to Amazon Facility for Scanning and Destruction for AI Training — SatelliteNetSec · 2026-08-17
- OpenAI & Oracle: Who Pays for the Power Bill? — aronchick · 2026-08-17
- Poll: Which Cloud Provider Do 2024+ Startups Use? — yenkel · 2026-08-17