LatentSeek Boosts LLM Reasoning in Latent Space via Policy Gradients
The paper "Seek in the Dark" introduces LatentSeek, a framework that enhances LLM reasoning via test-time instance-level adaptation in latent space. Instead of relying on chain-of-thought, it applies policy gradients directly to activations, guided by self-evaluation rewards.
2026-08-28 ~ 2026-08-28 · 2 related posts
- Paper Explores Policy Gradient over Activations for Reasoning — burny_tech · 2026-08-28
- LatentSeek Boosts LLM Reasoning via Test-Time Policy Gradients in Latent Space — burny_tech · 2026-08-28