GLM 5.2 runs at 35t/s via DwarfStar mixed RAM/VRAM inference

antirez · x · 2026-08-22

Antirez reports progress on the DGX Station: using DwarfStar mixed RAM/VRAM inference to run 4bit GLM 5.2 (500GB total weights). It achieves 35 tokens/s generation and 2000 tokens/s prefill (expected to hit 3k). The setup proves capable of running frontier open-weight models at agentic-usable speeds.

Original post →

More from Infra

Infra channel →