Local Inference Lab ships optimized Docker for RTX 6000 Pro / DGX Spark, loads GLM-5.3 in ~2 minutes

TheZachMueller · x · 2026-10-12

Local Inference Lab released optimized local-inference Docker images for NVIDIA RTX 6000 Pro and DGX Spark, with Karmic Kraken beta/stable channels and one-click configs for DeepSeek V4 Flash, GLM-5.3 (744B), MiMo V2.6, Qwen3.8 and more.

Highlights from the changelog: NVFP4-CSF checkpoints load in 2 minutes instead of 10-40; preempted requests reload prefixes from LMCache instead of recomputing; GLM-5.3 keeps decoding during long-prompt prefill; a fix for LMCache RAM-tier eviction on hybrid models; and GLM-5.3 prefills a 1M-token prompt at 92% GPU memory with decode context parallelism. Requires CUDA 13.4.1 / driver 615+; no DGX Spark image in this build.

Original post →

More from Infra

Infra channel →