phantom-kv Removes LLM Refusals via 18MB KV Injection

phantom-kv is an open-source project that removes LLM refusal behavior without touching model weights, by injecting an 18MB offline-trained KV tensor library into the model's KV cache at inference time, altering attention per request.

2026-09-22 ~ 2026-09-22 · 2 related posts

1 near-duplicate retellings: Anony6666