Laguna-S-2.1 runs on a 2020 PC with 8.3 GB VRAM and 52.2 GB RAM
crusaderky · reddit · 2026-07-23
Laguna-S-2.1 runs on a 6-year-old gaming PC with 8.3 GB VRAM and 52.2 GB RAM
A Reddit user reports that Laguna-S-2.1 runs on a 2020 desktop with an RTX 3080 10GB, Ryzen 9 3950X, and 64 GB RAM, though it is only practical overnight because memory headroom is extremely tight.
They say the model fits with IQ4XS, a 128k q8/q8 context, all dense layers on VRAM, and no DFlash. The setup reaches 120 t/s prefill and 10 t/s decode, using 8.3 GB VRAM and 52.2 GB host RAM. They also note that “thinking works out of the box” and expect memory pressure to improve once support lands in beellama and they can switch to kvarn.
They include their llamacpp CUDA configuration, showing settings such as ngl=99, n-cpu-moe=99, ctx-size=131072, flash-attn=on, and no-mmap=true.
More from Infra
- Alchemy beta expands to 204 AWS services, ECS/EKS, and Lambda workflows — samgoodwin89 · 2026-07-23
- Ten tactics to cut AI API bills by up to 90% without shrinking output — socialwithaayan · 2026-07-23
- Why AI pricing is really a memory-and-infrastructure problem, not just tokens — Kyrannio · 2026-07-23
- Open-source webfetch cuts agent search tokens by 87% and cost by 66% — Remote-Breadfruit204 · 2026-07-23
- NVIDIA launches Jetson Thor T2000 and T3000 for robotics and edge AI — NVIDIAAI · 2026-07-23
- AWS is said to ship 2.4 million Trainium3 accelerators this year — Beth_Kindig · 2026-07-23