Galahad Makes LLM Reading a One-Time Cost: 98.7% of Prompt Tokens Are Repeat Text

Sietse Schelpe · hf · 2026-10-01

Galahad is a byte-exact memory layer for vLLM, SGLang and llama.cpp that saves and reloads KV state for previously read text. With 98.7% of prompt tokens found to be repeat text, the system answers 100/100 fact-recall questions at 0.59-0.64s and 200-213J per question (vs 10/100 without it), with bit-identical logits after restart and support for all 30 models tested.

Original post →

More from Infra

Infra channel →