Running 276B MoE Models on <10GB RAM: Mference Hits ~3 tok/s

Blahblahblakha · reddit · 2026-08-06

A developer open-sourced Mference (built on Swift + Metal), successfully running the 276B parameter MoE model Inkling-Small (12B active) on consumer hardware with less than 10GB of memory.

Original post →

More from Infra

Infra channel →