What to run on a 96GB M3 Ultra? Qwen3.5-122B-A10B hits ~1000 tok/s locally
infieldmitt · reddit · 2026-09-17
A Reddit user asks what to run locally on a 96GB M3 Ultra Mac Studio, wanting both strong coding and emotionally intelligent writing. They complain current options are either 30B or way too big; running two Qwen3 27B instances to use the RAM feels futile since a faster MoE can preprocess quickly anyway. They just downloaded Qwen3.5-122B-A10B—up to 1000 tok/s and noticeably better prompt comprehension than smaller models—but worry the "3.5" naming means something inferior.
More from Infra
- Perovskite could lift solar efficiency ceiling from 30% to 45% — and give the US a shot against China — kyliebytes · 2026-09-17
- How mobile and specialization broke homogeneous compute into TPUs, NPUs, and more — blelbach · 2026-09-17
- After the x86 Monoculture: Software Will Suffer for Hardware's Fragmentation Again — blelbach · 2026-09-17
- Hardware veteran: low-precision gains nearly exhausted, true sparsity is AI's next 10x — blelbach · 2026-09-17
- Moore's Law in three eras: from free lunch (1970-2005) to software hell (2015-now) — blelbach · 2026-09-17
- Fed's first rate hike in 3 years raises the financing bar for debt-funded GPU clusters — rohanpaul_ai · 2026-09-17