LatentSeek Boosts LLM Reasoning via Test-Time Policy Gradients in Latent Space

burny_tech · x · 2026-08-28

The paper 'Seek in the Dark' introduces LatentSeek, a framework that enhances LLM reasoning through Test-Time Instance-level Adaptation (TTIA) within the model's latent space. Instead of operating in token space, it leverages policy gradients to iteratively update latent representations guided by self-generated reward signals. LatentSeek consistently outperforms strong baselines like Chain-of-Thought prompting and fine-tuning on benchmarks including GSM8K, MATH-500, and AIME2024.

Related event: LatentSeek Boosts LLM Reasoning in Latent Space via Policy Gradients(2 posts)→

Original post →

More from Research

Research channel →