Run a 744B MoE model on consumer hardware

Prompt Engineering · youtube · 2026-07-20

A demo of Colibri running GLM 5.2 on consumer hardware with about 25 GB of RAM.

The post highlights two key facts:

Because MoE routes only a subset of experts each step, the author argues that only roughly 11 GB of weights change token to token. The takeaway is that extremely large models can sometimes be made usable on modest hardware through routing and careful local execution.

Related event: Open-Source Colibri Runs 744B GLM-5.2 on Consumer Laptops(2 posts)→

Original post →

More from Infra

Infra channel →