Laguna S 2.1 on DGX Station claims 10 parallel 256K-context agents
max_paperclips · x · 2026-07-25
Laguna S 2.1 118B-A8B is profiled on a DGX Station with NVFP4 and FP8 KV cache, and the chart claims it can run 10 full-context agents with 256k context each.
- Model specs shown: 117.6B total, 8.5B active per token, 71.92 GB NVFP4.
- Serving setup: vLLM + FlashInfer, maxseqs=64, batchtokens=65,536, no swap, GPU utilization around 0.95.
- Reported capacity: 2,768,010 tokens FP8 KV cache and 10.56× 262,144-token sequences.
- Measured performance includes 202.42 tok/s per agent at 1K prompt/1-token decode, and 1,943.78 tok/s aggregate at 8K unique / concurrency 32.
More from Infra
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27
- MiniBot 2.40 adds xAI, HF Studio and vLLM support with inline media tools — Creative-Type9411 · 2026-07-27
- Apple smart glasses, Nvidia-SK AI data center deal, and Ctrip’s RMB 5.179 billion fine headline a tech roundup — APPSO · 2026-07-27
- DeepSeek funding rumor, EU AI transparency rules and OpenAI agent incident make a packed AI news roundup — 创业邦 · 2026-07-27
- QuixiCore argues native quantized kernels beat dequant-then-generic execution — QuixiAI · 2026-07-27