PKU PhD to present latent space reasoning optimization methods
青稞AI · wechat · 2026-08-18
Peking University PhD student Li Hengli will present two works on latent space reasoning for LLMs on Aug 22 via Qingke AI:
- LatentSeek: Uses self-reward and policy gradient for test-time optimization of latent representations, outperforming CoT on GSM8K and MATH.
- GradCuit: Solves credit assignment by allowing gradients to penetrate model circuits.
The talk covers the shift from likelihood-driven to value-driven reasoning and experimental results.
More from Research
- 3rd 3D HUMANS Workshop Returns at ECCV 2026 — dimadamen · 2026-08-18
- TogetherAI open-sources XoRL: 0 train-infer mismatch for large MoE RL training — PandaAshwinee · 2026-08-18
- SegDAC: Boosting Visual RL Generalization with Dynamic Object Tokens — GlenBerseth · 2026-08-18
- Study Compacts Context, Finds Prompt Caching Makes Summarization Obsolete — AI Engineer · 2026-08-18
- Georgia Tech's 2026 LLM Course: From MoE and Self-Play RL to Diffusion LMs — cocoweixu · 2026-08-18
- New Dataset: 35k Hugging Face Model Summaries Generated for $0.43 per 1k Rows — vanstriendaniel · 2026-08-18