Laptop engine streams a 35B model from SSD at 9.4 tok/s, beating GPT-OSS 20B
ImBadGuyInEveryStory · reddit · 2026-09-27
A poster's custom inference engine streams Ornith 1.5 35B (22GB) mostly from an SSD on a laptop with RTX 2050 4GB VRAM and 16GB RAM, hitting 9.4 tok/s — faster than GPT-OSS 20B's 6 tok/s in LM Studio. He asks whether pruned versions of top open MoE models (like Kimi K3) exist under 80GB at Q4 or above, ideally good for web dev, and how others have pruned older models like mimo v2.6.
More from Infra
- Three myths of hosted LLM inference: sticker prices, interchangeable endpoints, and self-hosting — TangeloOk9486 · 2026-09-27
- Random Attention: Salesforce and UIUC find random KV cache eviction rivals handcrafted signals — jiqizhixin · 2026-09-27
- World's fastest panel QR factorization on B200: how a GPU MODE contestant cracked chained dependencies — A_K_Nain · 2026-09-27
- TensorSharp open-source engine adds mixed document/image/video/audio inputs per request — fuzhongkai · 2026-09-27
- Pro-data center rally clashes with protesters as scholar defends AI infrastructure — neil_chilson · 2026-09-27
- Kafka in production: partition skew and rebalancing pauses bite hard — goyalshaliniuk · 2026-09-27