Local inference struggles with coding at 16 t/s, despite media use cases
natesiggard · x · 2026-08-26
The author lists viable local inference use cases like organizing a 25-year media library, Obsidian notes, and bookkeeping, but argues that 16 t/s is insufficient for 12+ hours of daily coding. Comments confirm that even on an M3 Ultra (512GB), GLM5.2 and Qwen 3.8 27B are capped around 16-30 t/s, sparking hope that the M5 will double this speed.
Related event: M3 Ultra Local LLM Tests Show Coding Still Out of Reach(2 posts)→
More from Infra
- Flexport cuts maintenance to zero with Anaconda orchestration — anacondainc · 2026-08-26
- Public favorability of data centers jumps 31% after learning about water recycling — BenBajarin · 2026-08-26
- Pipette tutorial: Benchmarking models on iPhone 13 Pro Max — helloiamleonie · 2026-08-26
- vLLM Sharded Weight Transfer Hits 7.53s for 1T Params Model — TheZachMueller · 2026-08-26
- Analyst: OpenAI and Anthropic validate the need for custom silicon — BenBajarin · 2026-08-26
- You don't need a Mac Studio for prompting — rudrank · 2026-08-26