Running a 35B Model With 131K Context on an 8GB Laptop: Full llama.cpp Config Shared

Aggravating-Push-207 · reddit · 2026-09-24

A user shares the full llama.cpp flags to run Cyber-Tiel-Coder-35B-A3B at 131K context on an 8GB VRAM 4060 laptop — CPU MoE offload, q40 KV cache, draft-MTP speculative decoding — hitting 35 tok/s prefill and 23 tok/s decode. The config was found by an OpenCode-driven overnight sweep.

Original post →

More from Infra

Infra channel →