Galahad Makes LLM Reading a One-Time Cost: 98.7% of Prompt Tokens Are Repeat Text
Sietse Schelpe · hf · 2026-10-01
Galahad is a byte-exact memory layer for vLLM, SGLang and llama.cpp that saves and reloads KV state for previously read text. With 98.7% of prompt tokens found to be repeat text, the system answers 100/100 fact-recall questions at 0.59-0.64s and 200-213J per question (vs 10/100 without it), with bit-identical logits after restart and support for all 30 models tested.
More from Infra
- DeepSeek releases TileKernels: dozens of LLM kernels near hardware limits, now on Ascend — rohanpaul_ai · 2026-10-01
- DeepSeek open-sources six Huawei Ascend projects, kernels hit 99.8% of chip peak — rohanpaul_ai · 2026-10-01
- Meta's Loop Scaling Laws: Sparsity Gives ~3x Active-Param Efficiency, Recurrence ~2x on Reasoning — facebook · 2026-10-01
- 87GB Qwen3.8 Flash Next runs at 120-150 t/s on a single RTX 5090 — CurieuxExplorer · 2026-10-01
- When evaluating GPU infra, "it passed testing" says little until you know what was tested — AccBalanced · 2026-10-01
- In the data center capital of the world, electricity rates declined from 2019-2024 — kevinnbass · 2026-10-01