Running Qwen3-27B on 16GB VRAM: struggling with pi.dev context compaction plugins

darksteelsteed · reddit · 2026-09-13

A user running NVFP4-quantized Qwen3.8-27B via llama.cpp on an RTX 5080 16GB (48k context, 12 t/s) reports pi.dev compacts at the wrong times, truncating responses. They tried npm:max-context (broken), pi-observational-memory (works partially at 0.75 threshold), and pi-blackhole (never compacts), and criticize pi.dev for keeping compaction settings separate from model settings—painful when switching between small-context local models. Full llama-server flags included.

Original post →

More from coding & agent

coding & agent channel →