Cherenkov engine hits 8-22 tok/s Qwen3.8-Flash-Next on a 32GB M4 MacBook Air
alfredr · reddit · 2026-09-11
- Developer released Cherenkov, an open-source inference engine for Apple Silicon that runs Qwen3.8-Flash-Next (Q4-ish mixed precision) at 8-22 tok/s using only 21GB of allocations — claimed record for memory-constrained inference of the model.
- Instead of loading the full model, it keeps a bounded working set of experts in unified memory, uses a one-layer lookahead to predict which experts are needed next, and prefetches them via SSD reads, with optional mixed-precision execution.
More from Infra
- k3 Report Section Confirms Millions of Concurrent Sandboxes in Its RL Training Run — stochasticchasm · 2026-09-11
- k3 RL Run Reportedly Used ~50M Sandboxes With Millions Concurrent, Checkpoints Merged Across Scaffolds — stochasticchasm · 2026-09-11
- Pentagon in talks to lend roughly $5 billion to AI cloud startup Fluidstack — vitaliychiley · 2026-09-11
- Eric Schmidt: AI may hit a money wall before a power wall — $1T capital needed — rohanpaul_ai · 2026-09-11
- SpaceX signs another AI compute deal: $1.11B per month, on track for $100B ARR — NinaDSchick · 2026-09-11
- Carmack: Jetson Thor's 128GB at 273GB/s is over-provisioned for real-time robotics — ID_AA_Carmack · 2026-09-11