What to run on a 96GB M3 Ultra? Qwen3.5-122B-A10B hits ~1000 tok/s locally

infieldmitt · reddit · 2026-09-17

A Reddit user asks what to run locally on a 96GB M3 Ultra Mac Studio, wanting both strong coding and emotionally intelligent writing. They complain current options are either 30B or way too big; running two Qwen3 27B instances to use the RAM feels futile since a faster MoE can preprocess quickly anyway. They just downloaded Qwen3.5-122B-A10B—up to 1000 tok/s and noticeably better prompt comprehension than smaller models—but worry the "3.5" naming means something inferior.

Original post →

More from Infra

Infra channel →