Paper Explores Policy Gradient over Activations for Reasoning
burny_tech · x · 2026-08-28
The post shares a paper titled "Seek in the Dark", proposing a method for performing reinforcement learning over activations in the latent space rather than over weights. This test-time instance-level policy gradient offers a new perspective for optimizing the reasoning process.
Related event: LatentSeek Boosts LLM Reasoning in Latent Space via Policy Gradients(2 posts)→
More from Research
- Using National Curricula to Build Culturally Grounded Data for LLMs — davlanade · 2026-08-28
- CDM speeds up reward-guided sampling in discrete diffusion by 50x with under 5% overhead — CatAstro_Piyush · 2026-08-28
- 7,400 trajectories analyzed: Claude Code and Codex carry ~25k ISL vs Terminus 2's 8k — zainhas · 2026-08-28
- Google's PPE Achieves Significant Gains in Epidemic and Food Security Prediction — ymatias · 2026-08-28
- Google Earth AI Unveils Planetary Prediction Engine for Geospatial Insights — ymatias · 2026-08-28
- METR report on Hugging Face attack hailed as first anthropology of posthuman civilization — anderssandberg · 2026-08-28