phantom-kv Removes LLM Refusals via 18MB KV Injection
phantom-kv is an open-source project that removes LLM refusal behavior without touching model weights, by injecting an 18MB offline-trained KV tensor library into the model's KV cache at inference time, altering attention per request.
2026-09-22 ~ 2026-09-22 · 2 related posts
- phantom-kv: uncensor LLMs per-request with an 18MB trained KV-cache, no weight edits — Anony6666 · 2026-09-22
1 near-duplicate retellings: Anony6666