Do models really 'pass messages through KV cache'? A skeptic pushes back
scaling01 · x · 2026-09-04
- Challenging the claim that models "pass messages through the KV cache far more complex than what they pass through token space."
- The author argues the KV cache is just the internal state given the context: recomputing all states without a cache yields the same result.
- So the cache isn't enabling any extra information passing — it only saves compute — and treating chain of thought as a proxy for the model's internal process may be misleading.
Related event: Looped transformer rumor sparks debate over CoT monitorability(6 posts)→
More from Models
- Benchmark saturation accelerates: ARC-AGI-3 falls in 5 months after ARC-AGI-1's 6 years — mmmbchang · 2026-09-04
- Sam Altman apologizes for messy Astra rollout, promises banked credit resets and broad API rollout soon — jxnlco · 2026-09-04
- GPT-6 went from shorthand for the far future to right on schedule, AI insider reflects — jeffintime · 2026-09-04
- GPT-6 'Astra' hype: builder McKay Wrigley calls it the biggest leap since GPT-4 — adrianscottcom · 2026-09-04
- Carina Hong launches Axiom to build a self-improving superintelligent AI mathematician — CatAstro_Piyush · 2026-09-04
- teortaxesTex: ability benchmarks are done, everything is environment — meta evals only — teortaxesTex · 2026-09-04