Muse Glimmer 30B Tested at 1M Context: Perfect Retrieval on Consumer Hardware
StartupTim · reddit · 2026-08-11
A developer successfully stretched Meta's day-old Muse Glimmer 30B context window from its trained 131K to 1M on a 2× NVIDIA DGX Spark cluster, passing 3/3 needle-in-haystack tests across all lengths.
- Hardware & Inference: Using llama.cpp with DFlash speculative decoding, single-node decode speed hit 36-38 tok/s (3x speedup). RPC split across both nodes worked but was 30% slower due to network hops.
- Long-context Theory: The author explains this perfect scaling by Muse's unique architecture: 13 global attention layers have NoPE (no positional encoding), while 39 sliding-window layers use RoPE only within a 2K window, making YaRN-stretching virtually harmless for long-range retrieval.
- Memory: Weights, drafter, vision module, and full 1M KV cache consume 60GB VRAM.
The author concludes it's a highly capable local agentic model and looks forward to vLLM support for RDMA multi-GPU deployment.
More from Infra
- Can a Single NVIDIA DGX Replace All Your AI Subscriptions? — jackedAJ · 2026-08-11
- MacBook + DGX Spark: Testing Heterogeneous Inference and KV Cache Shipping — HankYeomans · 2026-08-11
- RTX 3060 Tests: How GPU Memory Allocation Shifts LLM Execution Strategy — Abhishekcur · 2026-08-11
- CXMT 17nm DDR5 Yield Exceeds 90%, But US PC Makers Restrict Procurement — teortaxesTex · 2026-08-11
- Open-Source Models Cut Inference Costs 8x, Compute Becomes New Bottleneck — latticecut · 2026-08-11
- llama.cpp Adds CI Targets for ROCm 7.14, Boosting AMD GPU Inference — pmttyji · 2026-08-11