Beyond Token Generation: Training LLMs to Reason in Latent Working Memory

CShorten30 · x · 2026-08-04

Current LLM reasoning relies heavily on explicit token generation, causing high latency. The author introduces a new approach: training LLMs to reason within a latent working memory.

This method eliminates the visible reasoning trace, making Time To First Token (TTFT) equivalent to direct answering. An interactive blog post demonstrates the mechanism in action.

Original post →

More from Research

Research channel →