Running Qwen3-27B on 16GB VRAM: struggling with pi.dev context compaction plugins
darksteelsteed · reddit · 2026-09-13
A user running NVFP4-quantized Qwen3.8-27B via llama.cpp on an RTX 5080 16GB (48k context, 12 t/s) reports pi.dev compacts at the wrong times, truncating responses. They tried npm:max-context (broken), pi-observational-memory (works partially at 0.75 threshold), and pi-blackhole (never compacts), and criticize pi.dev for keeping compaction settings separate from model settings—painful when switching between small-context local models. Full llama-server flags included.
More from coding & agent
- Conductor hailed as best agentic dev environment: multi-agent workspaces, local sync — charlieholtz · 2026-09-13
- Ex-Cursor engineer claims $1.2M one-person company running 50 bots on a $200 plan — Roger_M_Taylor · 2026-09-13
- freeCodeCamp Releases Full Free OpenAI Codex Crash Course with Voice-Controlled Game Build — Roger_M_Taylor · 2026-09-13
- Where does multi-agent orchestration actually break in production? — RaraAvis27 · 2026-09-13
- SwiftUI Co-creator: "All Bugs Are Bridging Bugs" in Declarative-over-UIKit Stack — dotey · 2026-09-13
- Open-source reimagine-it MCP server turns pasted HTML into full redesigns — Thelastreddditor · 2026-09-13