GLM-5.2 runs locally on Dell Pro Max at 40 tokens/s, hinting at a new distillation pipeline
pcuenq · x · 2026-07-28
- Michael Dell shared a demo of GLM-5.2, a 753B-parameter model, running locally on a Dell Pro Max with GB300 at 40 tokens/s.
- The post argues that the exciting part is not local DIY inference for everyone, but that providers will serve models like this cheaply and fast.
- Expensive local and cloud setups will then distill these frontier models into much smaller ones that regular users can run.
- The broader point: frontier-scale intelligence is moving from the data center toward personal machines, and even models that are far too large for a $100K computer can still become teachers for future smaller systems.
More from Infra
- Jensen Huang backs open-source LLMs, but GPU costs remain the real barrier — pascalefung · 2026-07-28
- Google’s AlloyDB Omni demo runs fully offline with Gemma and TimesFM — rseroter · 2026-07-28
- Nvidia’s $750 billion AI deals reignite fears of circular financing — brainquantum · 2026-07-28
- SSI Signs $410M Compute Deal with Amazon, Bulk of Its Fundraising — TechCrunch AI · 2026-07-28
- Taiwan Detains Nvidia Employee in Widening AI Server Smuggling Probe — The Decoder · 2026-07-28
- RBC says a DIY AI replacement for Office could cost 11 times more over five years — TiernanRayTech · 2026-07-28