Laguna S 2.1 on DGX Station claims 10 parallel 256K-context agents
max_paperclips · x · 2026-07-25
Laguna S 2.1 118B-A8B is profiled on a DGX Station with NVFP4 and FP8 KV cache, and the chart claims it can run 10 full-context agents with 256k context each.
- Model specs shown: 117.6B total, 8.5B active per token, 71.92 GB NVFP4.
- Serving setup: vLLM + FlashInfer, maxseqs=64, batchtokens=65,536, no swap, GPU utilization around 0.95.
- Reported capacity: 2,768,010 tokens FP8 KV cache and 10.56× 262,144-token sequences.
- Measured performance includes 202.42 tok/s per agent at 1K prompt/1-token decode, and 1,943.78 tok/s aggregate at 8K unique / concurrency 32.
More from Infra
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11
- Local LLM server dilemma: 4x CMP-170HX (price up 53% in 20 days) vs Mac Studio M5 Ultra — rumboll · 2026-09-11
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11