744B Model Runs on 25GB RAM

techNmak · x · 2026-07-11

This post details an inference method to run a 744B parameter model on just 25GB RAM without a GPU. The trick isn't squeezing the whole model into memory, but ensuring most weights don't stay resident in RAM.

Core Concept

Practical Limitations

Unresolved Issues

Related event: 744B GLM-5.2 MoE Model Runs Locally on 25GB RAM(5 posts)→

Original post →

More from Infra

Infra channel →