GLM 5.2 runs at 35t/s via DwarfStar mixed RAM/VRAM inference
antirez · x · 2026-08-22
Antirez reports progress on the DGX Station: using DwarfStar mixed RAM/VRAM inference to run 4bit GLM 5.2 (500GB total weights). It achieves 35 tokens/s generation and 2000 tokens/s prefill (expected to hit 3k). The setup proves capable of running frontier open-weight models at agentic-usable speeds.
More from Infra
- Matryoshka Framework: Train Model Suites 36% Cheaper with Nested Architecture — TheTuringPost · 2026-08-22
- Stanford CS336 wraps up with deep dive into GPU programming and frontier inference — stanfordnlp · 2026-08-22
- How Pi handles context compaction for long coding sessions — bibryam · 2026-08-22
- Gwangyang Steel Works reveals the poverty of the data center energy debate — AndyMasley · 2026-08-22
- Benchmark Shows Unified Memory Significantly Boosts Local LLM VRAM Efficiency — Pablo_the_brave · 2026-08-22
- OpenAI Acquires Backend Startup Instant to Bolster AI App Infrastructure — testingcatalog · 2026-08-22