Paper Explores Policy Gradient over Activations for Reasoning

burny_tech · x · 2026-08-28

The post shares a paper titled "Seek in the Dark", proposing a method for performing reinforcement learning over activations in the latent space rather than over weights. This test-time instance-level policy gradient offers a new perspective for optimizing the reasoning process.

Related event: LatentSeek Boosts LLM Reasoning in Latent Space via Policy Gradients(2 posts)→

Original post →

More from Research

Research channel →