Local Inference Lab ships optimized Docker for RTX 6000 Pro / DGX Spark, loads GLM-5.3 in ~2 minutes
TheZachMueller · x · 2026-10-12
Local Inference Lab released optimized local-inference Docker images for NVIDIA RTX 6000 Pro and DGX Spark, with Karmic Kraken beta/stable channels and one-click configs for DeepSeek V4 Flash, GLM-5.3 (744B), MiMo V2.6, Qwen3.8 and more.
Highlights from the changelog: NVFP4-CSF checkpoints load in 2 minutes instead of 10-40; preempted requests reload prefixes from LMCache instead of recomputing; GLM-5.3 keeps decoding during long-prompt prefill; a fix for LMCache RAM-tier eviction on hybrid models; and GLM-5.3 prefills a 1M-token prompt at 92% GPU memory with decode context parallelism. Requires CUDA 13.4.1 / driver 615+; no DGX Spark image in this build.
More from Infra
- Running a 456GB model on 192GB VRAM: offloaded inference hits 60-125 tok/s with 1M context — HankYeomans · 2026-10-12
- VitalOps launches agentic inference optimization, 2.6x median speedup — abhijithneil · 2026-10-12
- Running Qwen3.8 Flash-Next locally on AMD 7900 XTX at 500k context, 105-160 tok/s — human_in_the_looop · 2026-10-12
- Siemens Brings Nvidia Omniverse into Digital Twin Composer to Pave the Way for Physical AI — RevLebaredian · 2026-10-12
- OpenRouter token traffic explodes from 2T to 379T/month, open-weight models at 75% — Beth_Kindig · 2026-10-12
- Struggling to match Strata speeds running Qwen3.8-Flash-Next in vanilla llama.cpp on 2x RTX 3060 — Dreeew84 · 2026-10-12