Running a 35B Model With 131K Context on an 8GB Laptop: Full llama.cpp Config Shared
Aggravating-Push-207 · reddit · 2026-09-24
A user shares the full llama.cpp flags to run Cyber-Tiel-Coder-35B-A3B at 131K context on an 8GB VRAM 4060 laptop — CPU MoE offload, q40 KV cache, draft-MTP speculative decoding — hitting 35 tok/s prefill and 23 tok/s decode. The config was found by an OpenCode-driven overnight sweep.
More from Infra
- Why the Semiconductor Industry Is Betting on Optical Interconnects for AI — BenBajarin · 2026-09-24
- Qualcomm Officially Brings Linux to Snapdragon X2 Chips, Debian Coming End of This Year — tomwarren · 2026-09-24
- PyTorchCon 2026 Poster to Show Warm-Start Autotuning for Helion GPU Kernel DSL — PyTorch · 2026-09-24
- Qualcomm Announces Linux Support for Snapdragon X2 Elite — karlfreund · 2026-09-24
- MediaTek's consensus growth seen at 65-85% on TPU demand, analyst says — BenBajarin · 2026-09-24
- Company From 'Gas to Gigawatts' Expert Interview Lands on UBS Most Preferred List — BenBajarin · 2026-09-24