KV Cache as an Information Bottleneck Between Tokens: Austin Huang's Idea
austinvhuang · x · 2026-09-02
Austin Huang resurfaces his 2024 note: now that everyone treats the KV cache as transformer state, the next step is thinking of the KV cache as an information bottleneck between tokens — a framing relevant to attention-state compression trade-offs.
More from Research
- A real neural net trained inside the EVM: 200K gas inference, ~340K gas per training step — jarrodwatts · 2026-09-02
- Cosmos Institute Funds 80 Projects Building AI for Human Autonomy and Truth-Seeking — lawhsw · 2026-09-02
- Every major AI model fails at polytonic Ancient Greek — and RLHF makes it worse — vasilisvj · 2026-09-02
- Bridgewater fine-tuned Qwen3-235B to 84.7%, beating frontier models at 1/14 the cost — anacondainc · 2026-09-02
- NextLat's latent variable-length speculative decoding speeds inference up to 3.3x — morgymcg · 2026-09-02
- Papers with Code spotlights Looped Transformer: recurrent depth decoupled from parameters — NielsRogge · 2026-09-02